IntelliAuto
Vehicle management. AutoMind AI, ML Kit OCR, offline encrypted storage.
SYS.01 Systems engineering · local AI · product
Enterprise Linux operations, LLM inference on Intel Arc Pro B70, and production Android apps — documented with the trade-offs left in.
Lifecycle, automation, hardening, HA and the operational work that keeps platforms predictable.
Local AI Lab
Intel Arc Pro B70 32GB running llama.cpp SYCL — every number below is hardware-measured, not estimated.
Production Apps
Two privacy-minded Android products built around local-first value and resilient cloud boundaries.
Vehicle management. AutoMind AI, ML Kit OCR, offline encrypted storage.
Finance tracker. Offline-first, cloud sync, budget projections, spending summaries.
03 Educational Series
An engineer's guide to Local AI. Demystifying Transformer architectures, KV cache quantization, modern attention and MoE routing.
Recent Field Notes
The MTP draft head is recurrent, so a single-layer MTP checkpoint emits N speculative tokens per step. The spec-N curve (N=1/2/4) is the real story: N=4 wins throughput at 204.6 t/s single-stream decode (+41% vs the community 145 claim) even though per-token acceptance falls with N (92.6% → 84.2% → 72.7% → 68.9% by draft position). N=1 reaches 97.1% acceptance — but costs 30% throughput. The right objective for a single-layer draft head is throughput, not acceptance %. A paired, alternating power A/B also killed the 'boost prefill to 230W' hypothesis: 150W vs 230W prefill is flat at ±0.2%.