IntelliAuto
Vehicle management. AutoMind AI, ML Kit OCR, offline encrypted storage.
SYS.01 Systems engineering · local AI · product
Enterprise Linux operations, LLM inference on Intel Arc Pro B70, and production Android apps — documented with the trade-offs left in.
Lifecycle, automation, hardening, HA and the operational work that keeps platforms predictable.
Local AI Lab
Intel Arc Pro B70 32GB. Each card is one engine, one model, one cell — not a blended leaderboard.
Production Apps
Two privacy-minded Android products built around local-first value and resilient cloud boundaries.
Vehicle management. AutoMind AI, ML Kit OCR, offline encrypted storage.
Finance tracker. Offline-first, cloud sync, budget projections, spending summaries.
03 Educational Series
An engineer's guide to Local AI. Demystifying Transformer architectures, KV cache quantization, modern attention and MoE routing.
vLLM XPU · 3B-active LatentMoE 186.6 t/s C1 with DFlash, not native MTP
Graphs plus a BF16 DFlash draft. Native MTP still accepts 0%. Speed card is 16K n=5; capacity completed 119,904 tokens. Not 128K.
llama.cpp SYCL · dense multimodal 26.8 t/s at 128K — vLLM is not the public path
DFlash n2 wins the GGUF sweep. The 29.2 number is 64K. INT4 vLLM exists only as a local overlay.
llama.cpp SYCL · dense 27B + MoE 35B The living recipe, then the 27B MTP card
Qwen numbers stay on their own posts. Do not mix them with Nemotron client-TTFT or DFlash cells.
KV · power · context ceilings Why q8_0 K + q4_1 V, and what 32 GB actually holds
Methodology posts. These are not model speed cards.
Brand logos are trademarks of their respective owners; shown for identification.
Rest of posts