SYS.01 Systems engineering · local AI · product

Linux infra, local AI, shipped apps.

Enterprise Linux operations, LLM inference on Intel Arc Pro B70, and production Android apps — documented with the trade-offs left in.

system_map://sergiiob.dev

interactive
01 / 04 PLATFORM
Enterprise Linux

Lifecycle, automation, hardening, HA and the operational work that keeps platforms predictable.

Select a node to inspect the system
01

Local AI Lab

One GPU, measured to the token.

Intel Arc Pro B70 32GB. Each card is one engine, one model, one cell — not a blended leaderboard.

02

Production Apps

Useful software, not AI wrappers.

Two privacy-minded Android products built around local-first value and resilient cloud boundaries.

ANDROID / VEHICLE OPS Live on Play Store
IntelliAuto app icon

IntelliAuto

Vehicle management. AutoMind AI, ML Kit OCR, offline encrypted storage.

KotlinFirebaseOffline-firstML Kit
ANDROID / FINANCE OPS Live on Play Store
IntelliFlow app icon

IntelliFlow

Finance tracker. Offline-first, cloud sync, budget projections, spending summaries.

KotlinFirebasePrivacy-firstForecasting

03 Educational Series

The LLM Infrastructure Handbook

An engineer's guide to Local AI. Demystifying Transformer architectures, KV cache quantization, modern attention and MoE routing.

Read the Handbook
04

Local LLM Recipes

19 B70 index

Brand logos are trademarks of their respective owners; shown for identification.

05

Rest of posts

More field notes

01

Qwen3.8-27B FP8 on Two Intel Arc Pro B70s: TP2, Graphs, and Context

02

Qwen3.8-27B on Windows 11 + Arc Pro B70: the 19 August upgrade

03

Phase 1: The vLLM Question on Intel Arc Pro B70 (MXFP4 Native Test)

04

118B MoE on a single 32GB GPU: Laguna S 2.1 partial expert offload

05

Arc Pro B70 clean suite: Gemma 4 31B MTP, MoE prefill, and Grok tools

06

Grok Build CLI with local models on llama-server (Arc Pro B70)

07

The Reality of Edge AI Research: Why TurboQuant on Intel Arc SYCL Failed (For Now)

08

Breaking the 67 tok/s Barrier: Optimizing Intel Arc Pro B70 for High-Concurrency MoE Inference

09

MTP-4 Speculative Decoding Power Scaling and Benchmark Methodology Fix on Intel Arc B70

10

Intel Arc Pro B70 32GB: Running Qwen3.6-35B on llama.cpp SYCL

11

Optimizing DeepSeek KV Cache for Serverless AI Pipelines

12

RX 7800 XT 16GB: Running 35B MoE at 128K Context with llama.cpp + ROCm

13

Git Branch Splitting: Untangling Mixed Feature Branches

14

14 Models Benchmarked on RK3588: The Definitive CPU vs NPU Ranking

15

llamacpp-workbench: Remote llama.cpp Control and REAP Model Serving on RK3588

16

Qwen3.5 on RK3588 with llama.cpp: Real Benchmarks from a Radxa ROCK 5B+

17

GPU VRAM, CPU Offload, and llama.cpp: The Real Performance Cliff

18

Implementing Google's TurboQuant: Hybrid KV Cache for Edge LLM Deployment

19

The Architecture of Speed: Real-Time Telemetry and Generative AI in 2026 Motorsport

20

RK3588 NPU Router Architecture: What Actually Runs, What Wins, and Why

21

The Comprehensive Linux Engineer Command List

22

Architecting for Stability: Replacing Legacy IP-Binding with Layer 7 Proxy Routing

23

Azure Provisioned Throughput: When Fixed Costs Beat Pay-Per-Token

24

Ansible Vault, Python, and Molecule Snippets

25

Daily Linux Ops Command Cheatsheet

26

Training Custom AI Models for Insurance Document Processing

27

DNS Migration Strategy for Zero-Downtime System Replacement

28

Edge LLM Optimization: Memory Bandwidth and Context Management

29

Enterprise Certificate Lifecycle Management with Ansible

30

IntelliAuto: AI-Powered Automotive Assistant with Secure Monetization

31

IntelliFlow: Building a Production-Ready Finance App with AI

32

Linux-Active Directory Integration: Access Control, SSO, and Troubleshooting

33

LVM Operations: Expand, Shrink, and Migrate Volumes

34

Automating NFS Share Management at Scale

35

Server Provisioning Playbook: From VM Request to Production

36

Security Layering for Edge AI APIs: Encryption, Rate Limits, Validation, and Monitoring

37

Automating Firebase Deployments: Multi-Account Routing and Discord Notifications

38

RK3588 LLM Performance: NPU vs CPU in a Discord Agent

39

Stretched Networks and Leaf-Spine Architecture

40

Automating AD Computer Object Deletion on Linux Decommission

41

Modernizing Android UX: High Refresh Rates & App Shortcuts

42

Silent Software Installations with Ansible

43

Building Custom Ansible Execution Environments

44

PostgreSQL WAL Archiving with SELinux Considerations

45

Apache as a Reverse Proxy: Ansible Deployment Pattern

46

Shipping My First Android App: IntelliFlow

47

Building Golden Images with Packer and StackGuardian

48

Securing and Scaling AI Context in an Automotive Assistant

49

Orchestrating Patching Waves for Enterprise Linux

50

Testing Ansible Roles with Molecule and Docker

51

Building a Multilingual AI Backend for Part Recognition

52

Slashing LLM API Costs with System Prompt Caching

53

Managing Linux Users and Groups with Ansible

54

Bash Script for User Permission Audits

55

Implementing the Outbox Pattern for Offline-First Sync

56

Tracking Required Reboots with RHEL Tracer

57

Essential Red Hat Linux Administrator Commands

58

Safely Resolving Git Merge Conflicts

59

Infrastructure as Code: Structuring Ansible Repositories

68 Full chronological archive