SYS.01 Systems engineering · local AI · product

Linux infra, local AI, shipped apps.

Enterprise Linux operations, LLM inference on Intel Arc Pro B70, and production Android apps — documented with the trade-offs left in.

system_map://sergiiob.dev

interactive
01 / 04 PLATFORM
Enterprise Linux

Lifecycle, automation, hardening, HA and the operational work that keeps platforms predictable.

Select a node to inspect the system
01

Local AI Lab

One GPU, measured to the token.

Intel Arc Pro B70 32GB. Each card is one engine, one model, one cell — not a blended leaderboard.

02

Production Apps

Useful software, not AI wrappers.

Two privacy-minded Android products built around local-first value and resilient cloud boundaries.

ANDROID / VEHICLE OPS Live on Play Store
IntelliAuto app icon

IntelliAuto

Vehicle management. AutoMind AI, ML Kit OCR, offline encrypted storage.

KotlinFirebaseOffline-firstML Kit
ANDROID / FINANCE OPS Live on Play Store
IntelliFlow app icon

IntelliFlow

Finance tracker. Offline-first, cloud sync, budget projections, spending summaries.

KotlinFirebasePrivacy-firstForecasting

03 Educational Series

The LLM Infrastructure Handbook

An engineer's guide to Local AI. Demystifying Transformer architectures, KV cache quantization, modern attention and MoE routing.

Read the Handbook
04

Local LLM Recipes

19 B70 index

Brand logos are trademarks of their respective owners; shown for identification.

05

Rest of posts

More field notes

01

Renewing Red Hat Satellite's HTTPS Certificate (katello-certs-check + satellite-installer)

02

Qwen3.8-27B FP8 on Two Intel Arc Pro B70s: TP2, Graphs, and Context

03

Qwen3.8-27B on Windows 11 + Arc Pro B70: the 19 August upgrade

04

Phase 1: The vLLM Question on Intel Arc Pro B70 (MXFP4 Native Test)

05

118B MoE on a single 32GB GPU: Laguna S 2.1 partial expert offload

06

Arc Pro B70 clean suite: Gemma 4 31B MTP, MoE prefill, and Grok tools

07

Grok Build CLI with local models on llama-server (Arc Pro B70)

08

The Reality of Edge AI Research: Why TurboQuant on Intel Arc SYCL Failed (For Now)

09

Breaking the 67 tok/s Barrier: Optimizing Intel Arc Pro B70 for High-Concurrency MoE Inference

10

MTP-4 Speculative Decoding Power Scaling and Benchmark Methodology Fix on Intel Arc B70

11

Intel Arc Pro B70 32GB: Running Qwen3.6-35B on llama.cpp SYCL

12

Optimizing DeepSeek KV Cache for Serverless AI Pipelines

13

RX 7800 XT 16GB: Running 35B MoE at 128K Context with llama.cpp + ROCm

14

Git Branch Splitting: Untangling Mixed Feature Branches

15

14 Models Benchmarked on RK3588: The Definitive CPU vs NPU Ranking

16

llamacpp-workbench: Remote llama.cpp Control and REAP Model Serving on RK3588

17

Qwen3.5 on RK3588 with llama.cpp: Real Benchmarks from a Radxa ROCK 5B+

18

GPU VRAM, CPU Offload, and llama.cpp: The Real Performance Cliff

19

Implementing Google's TurboQuant: Hybrid KV Cache for Edge LLM Deployment

20

The Architecture of Speed: Real-Time Telemetry and Generative AI in 2026 Motorsport

21

RK3588 NPU Router Architecture: What Actually Runs, What Wins, and Why

22

The Comprehensive Linux Engineer Command List

23

Architecting for Stability: Replacing Legacy IP-Binding with Layer 7 Proxy Routing

24

Azure Provisioned Throughput: When Fixed Costs Beat Pay-Per-Token

25

Ansible Vault, Python, and Molecule Snippets

26

Daily Linux Ops Command Cheatsheet

27

Training Custom AI Models for Insurance Document Processing

28

DNS Migration Strategy for Zero-Downtime System Replacement

29

Edge LLM Optimization: Memory Bandwidth and Context Management

30

Enterprise Certificate Lifecycle Management with Ansible

31

IntelliAuto: AI-Powered Automotive Assistant with Secure Monetization

32

IntelliFlow: Building a Production-Ready Finance App with AI

33

Linux-Active Directory Integration: Access Control, SSO, and Troubleshooting

34

LVM Operations: Expand, Shrink, and Migrate Volumes

35

Automating NFS Share Management at Scale

36

Server Provisioning Playbook: From VM Request to Production

37

Security Layering for Edge AI APIs: Encryption, Rate Limits, Validation, and Monitoring

38

Automating Firebase Deployments: Multi-Account Routing and Discord Notifications

39

RK3588 LLM Performance: NPU vs CPU in a Discord Agent

40

Stretched Networks and Leaf-Spine Architecture

41

Automating AD Computer Object Deletion on Linux Decommission

42

Modernizing Android UX: High Refresh Rates & App Shortcuts

43

Silent Software Installations with Ansible

44

Building Custom Ansible Execution Environments

45

PostgreSQL WAL Archiving with SELinux Considerations

46

Apache as a Reverse Proxy: Ansible Deployment Pattern

47

Shipping My First Android App: IntelliFlow

48

Building Golden Images with Packer and StackGuardian

49

Securing and Scaling AI Context in an Automotive Assistant

50

Orchestrating Patching Waves for Enterprise Linux

51

Testing Ansible Roles with Molecule and Docker

52

Building a Multilingual AI Backend for Part Recognition

53

Slashing LLM API Costs with System Prompt Caching

54

Managing Linux Users and Groups with Ansible

55

Bash Script for User Permission Audits

56

Implementing the Outbox Pattern for Offline-First Sync

57

Tracking Required Reboots with RHEL Tracer

58

Essential Red Hat Linux Administrator Commands

59

Safely Resolving Git Merge Conflicts

60

Infrastructure as Code: Structuring Ansible Repositories

69 Full chronological archive