<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Shreyas on Systems]]></title><description><![CDATA[Thought my local 14B AI agent was eating, but it was secretly dropping parameters and timing out. How to actually fix local models on a 16GB GPU.]]></description><link>https://shreyas-systems.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6ab2fac24987b06129891efa/e8e64d39-f6aa-4e4d-a4a5-f6e103dcb5f5.png</url><title>Shreyas on Systems</title><link>https://shreyas-systems.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Mon, 05 Oct 2026 09:53:13 GMT</lastBuildDate><atom:link href="https://shreyas-systems.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[I Spent 3 Weeks Debugging a Local LLM Agent on Consumer Hardware. Here’s Everything That Broke.]]></title><description><![CDATA[I recently set out to build an air-gapped, local-first autonomous agent running on consumer hardware. My constraints were strict: a single 16GB VRAM graphics card (an AMD Radeon RX 7800 XT), an Ollama]]></description><link>https://shreyas-systems.hashnode.dev/debugging-local-llm-agents</link><guid isPermaLink="true">https://shreyas-systems.hashnode.dev/debugging-local-llm-agents</guid><category><![CDATA[AI]]></category><category><![CDATA[Python]]></category><category><![CDATA[Open Source]]></category><category><![CDATA[Machine Learning]]></category><category><![CDATA[architecture]]></category><category><![CDATA[ollama]]></category><dc:creator><![CDATA[Shreyas Datta]]></dc:creator><pubDate>Thu, 24 Sep 2026 23:35:44 GMT</pubDate><enclosure url="https://cdn.hashnode.com/uploads/covers/6ab2fac24987b06129891efa/bbe92af2-fc67-43b8-a0ab-1265e65b207b.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>I recently set out to build an air-gapped, local-first autonomous agent running on consumer hardware. My constraints were strict: a single 16GB VRAM graphics card (an AMD Radeon RX 7800 XT), an Ollama backend running <code>qwen2.5-coder:14b</code>, pure Python, and zero bloated frameworks like LangChain or AutoGen.</p>
<p>The architectural philosophy was straightforward: treat the Large Language Model as a stateless reasoning coprocessor—an Arithmetic Logic Unit (ALU)—while the host Python application acts as the operating system, owning routing, memory, and execution boundaries.</p>
<p>After weeks of development, I spun up an automated 40-case regression suite, hit enter, and watched the terminal print a perfect result:</p>
<pre><code class="language-plaintext">============================================================
IDD EXECUTION RECEIPT | LIVE REGRESSION SUITE
Model: qwen2.5-coder:14b-instruct-q4_K_M
Cases: 40 | PASS: 40 | FAIL: 0
============================================================
</code></pre>
<blockquote>
<p>Forty tests executed, forty passed.</p>
</blockquote>
<p>However, inspecting the host terminal telemetry and raw Ollama server logs revealed that the system was actually collapsing under the hood. The green checkmarks were an illusion.</p>
<hr />
<img src="https://cdn.hashnode.com/uploads/covers/6ab2fac24987b06129891efa/5ce23af7-f732-42de-b457-e58443d127e0.png" alt="" style="display:block;margin:0 auto" />

<h2>I. The Illusion of "Assertion Slack"</h2>
<p>In traditional software, if \(2 + 2 = 5\), your test crashes and fails immediately. In agentic software, a test suite only fails if an unhandled exception surfaces.</p>
<p>My regression harness checked high-level status strings:</p>
<ul>
<li><p>Did the turn return <code>status == "completed"</code>?</p>
</li>
<li><p>Did the security interceptor report <code>status == "confirmation_required"</code>?</p>
</li>
</ul>
<p>The agent was returning those exact strings. But buried deep in the execution traces was a repeating error:</p>
<p>Plaintext</p>
<pre><code class="language-plaintext">[Action] Ephemeral Extraction: calculator with {}
[Action] Ephemeral Extraction: write_file with {}
</code></pre>
<p>Every single tool call was passing an empty argument dictionary <code>{}</code>. When the agent attempted arithmetic, it extracted no numbers and hallucinated a guess from its parametric memory. When it tried to write a file, it extracted no path and no content.</p>
<p>The test passed because the Human-In-The-Loop safety breaker caught the action and paused for approval. Had a user approved the action, the script would have crashed instantly due to missing arguments.</p>
<hr />
<h2>II. The Two Silent Killers of Local 14B Models</h2>
<p>Running open-weight models locally on Ollama or <code>llama.cpp</code> introduces low-level system failures that cloud APIs abstract away.</p>
<h3>A. Failure 1: The Leaky Wire Format (ACI Serialization Mismatch)</h3>
<p>Client libraries expect tool calls to arrive neatly inside a structured JSON array (<code>response.choices[0].message.tool_calls</code>).</p>
<p>Ollama detects tool calls by parsing raw model tokens against an internal Go-based template engine. Open-weight models like Qwen 2.5-Coder often emit tool calls wrapped in inline text tags (such as <code>&lt;tool_call&gt;{"name": "...", "arguments": {...}}&lt;/tool_call&gt;</code>) directly inside the conversational body (<code>message.content</code>).</p>
<p>Because the Python engine strictly looked for <code>message.tool_calls</code>, it found an empty list, assumed the model had generated no arguments, and defaulted to <code>{}</code>. The parameters existed in the raw stream, but the interface missed them completely.</p>
<h3>B. Failure 2: The Sampler Freeze (GBNF Logit Locks)</h3>
<p>To force the model to output strict JSON during intent routing, I applied constrained grammar decoding (<code>format={"type": "object", ...}</code>).</p>
<p>Suddenly, simple conversational questions like <em>"What is software coupling?"</em> began stalling for 14 to 40 seconds before triggering socket timeouts.</p>
<p>When a chat-tuned model encounters a conceptual prompt, its training predisposes it to converse first (e.g., <em>"Coupling refers to..."</em>) before producing structural brackets. However, the constrained decoding engine sets the probability of all non-JSON tokens to zero.</p>
<p>The model attempts to output natural language, the sampler rejects the tokens, and the engine enters a low-probability evaluation loop. Compute cycles burn evaluating rejected logits until the HTTP socket terminates.</p>
<hr />
<h2>III. The Architecture War: Regex vs. Pure LLM Probing</h2>
<p>To fix this, the agent needed an intent router to determine whether a query required a tool or a direct conversational response.</p>
<pre><code class="language-plaintext">                           Incoming Query
                                 │
                                 ▼
             ┌───────────────────────────────────────┐
             │ Stage 1: Positive Regex Fast-Path     │
             │         (Latency &lt; 1 ms)              │
             └───────────────────┬───────────────────┘
                                 │
              Matched? ──────────┼──────────┐
              (e.g., "749 x 192")│          │
                                 │          ▼
                                NO       YES ──► Execute Tool Directly (&lt;1 ms)
                                 │               (Bypasses GPU Prefill)
                                 ▼
             ┌───────────────────────────────────────┐
             │ Stage 2: Negative-Anchored LLM Probe  │
             │         (Latency ~150-250 ms)         │
             └───────────────────────────────────────┘
</code></pre>
<h3>I initially tested two common design approaches:</h3>
<ul>
<li><p><strong>Attempt 1: Brittle Regex Filters.</strong> Simple regex matching digits and operators (<code>\d+ * \d+</code>) resolves calculations in <code>&lt; 1ms</code>. However, regex lacks semantic awareness. If a user asks, <em>"Review this C++ buffer initialization:</em> <code>capacity = 1024 * 16;</code><em>"</em>, the pattern detects <code>1024 * 16</code>, hijacks the query, runs the calculator, and completely ignores the code review request.</p>
</li>
<li><p><strong>Attempt 2: Pure LLM Intent Probing.</strong> Passing every query through a classification prompt avoids regex collisions. However, to resolve conversational pronouns (e.g., <em>"Now multiply that by 8"</em>), the probe must ingest recent chat history. Ingesting dynamic history invalidates the GPU's Key-Value (KV) cache prefix, forcing a full prefill computation that turns a 1ms operation into a 300ms to 800ms bottleneck.</p>
</li>
</ul>
<hr />
<h2>IV. The Solution: Guard-Gated Two-Phase Decoupled Dispatch</h2>
<p>The finalized architecture separates conversational generation, tool extraction, and deterministic execution into distinct stages:</p>
<h3>Layer 1: The Static System Anchor (100% KV-Cache Reuse)</h3>
<p>The primary system prompt remains fixed at Position 0 across the entire session with <code>tools=None</code> enforced. By omitting tool definitions from conversational passes, the model experiences zero "tool magnetism" (the tendency for small models to invoke tools randomly on greetings).</p>
<h3>Layer 2: Tiered Intent Routing</h3>
<ol>
<li><p><strong>Python Code Guard:</strong> The query is scanned for programming keywords (<code>class</code>, <code>def</code>, <code>struct</code>, <code>uint32_t</code>). If detected, regex matching is disabled to prevent code snippets from triggering tools.</p>
</li>
<li><p><strong>Positive Syntax Accelerator:</strong> Explicit arithmetic and clock queries match a positive regex and execute in <code>&lt; 1ms</code> on the CPU, bypassing the GPU.</p>
</li>
<li><p><strong>Negative-Anchored Semantic Fallback:</strong> Unstructured queries fall through to an isolated LLM probe. The system prompt contains explicit negative constraints (<em>"DO NOT answer the prompt. DO NOT explain concepts. Output ONLY JSON."</em>), preventing token-sampler stalls.</p>
</li>
</ol>
<h3>Layer 3: Ephemeral Dual-Path Parameter Extraction</h3>
<p>When a tool is required, an isolated single-turn request runs containing only the target tool schema.</p>
<ul>
<li><p><strong>Path A:</strong> Inspect <code>message.tool_calls</code> for standard JSON arguments.</p>
</li>
<li><p><strong>Path B:</strong> If empty, an extraction fallback scans <code>message.content</code> using substring slicing to extract raw <code>&lt;tool_call&gt;</code> wrappers or embedded JSON objects.</p>
</li>
</ul>
<h3>Layer 4: Tail-Injected Synthesis</h3>
<p>Verified tool results are appended exclusively to the end of the user prompt inside <code>&lt;observation&gt;</code> XML blocks. The model summarizes the output in conversational mode with <code>tools=None</code>, maintaining full context without cache eviction.</p>
<hr />
<h2>V. Telemetry Validation: Before vs. After</h2>
<p>Rerunning the regression harness on the decoupled architecture demonstrated immediate performance gains across the stack:</p>
<table style="min-width:292px"><colgroup><col style="min-width:25px"></col><col style="min-width:25px"></col><col style="width:167px"></col><col style="min-width:25px"></col><col style="min-width:25px"></col><col style="min-width:25px"></col></colgroup><tbody><tr><td><p>Performance <strong>Metric</strong></p></td><td><p><strong>Baseline Implementation</strong></p></td><td><p><strong>Decoupled Dispatch Pipeline</strong></p></td><td><p></p></td><td><p></p></td><td><p></p></td></tr><tr><td><p><strong>Tool Argument Extraction Failure</strong></p></td><td><p>100% (Returned <code>{}</code>)</p></td><td><p><strong>0.0% (Exact Parameter Capture)</strong></p></td><td><p></p></td><td><p></p></td><td><p></p></td></tr><tr><td><p><strong>Arithmetic Precision (\(749 \times 192\))</strong></p></td><td><p>Failed ($144,368$ via memory guess)</p></td><td><p><strong>Passed ($143,808$ via execution)</strong></p></td><td><p></p></td><td><p></p></td><td><p></p></td></tr><tr><td><p><strong>Router Latency (Conversational)</strong></p></td><td><p>\(14,684\text{ ms} - 39,767\text{ ms}\) (Stalls)</p></td><td><p><strong>\(118\text{ ms} - 195\text{ ms}\)</strong></p></td><td><p></p></td><td><p></p></td><td><p></p></td></tr><tr><td><p><strong>Position 0 Cache Hit Rate</strong></p></td><td><p>0% (Prefix thrashing)</p></td><td><p><strong>100% (RadixAttention retained)</strong></p></td><td><p></p></td><td><p></p></td><td><p></p></td></tr><tr><td><p><strong>Prompt Evaluation Duration</strong></p></td><td><p>\(484.60\text{ ms} - 875.24\text{ ms}\)</p></td><td><p><strong>\(57.47\text{ ms} - 75.35\text{ ms}\)</strong></p></td><td><p></p></td><td><p></p></td><td><p></p></td></tr><tr><td><p><strong>VRAM Footprint Allocation</strong></p></td><td><p>13.9 GB (Near overflow)</p></td><td><p><strong>8.8 GB / 16.0 GB (Stable)</strong></p></td><td><p></p></td><td><p></p></td><td><p></p></td></tr></tbody></table>

<h2>Hard-Earned Lessons for Local Agent Builders</h2>
<ol>
<li><p><strong>Do not trust high-level test receipts.</strong> Always write assertions against ground-truth payloads, non-empty argument dictionaries, and exact tool outputs.</p>
</li>
<li><p><strong>The LLM is an ALU, not an Operating System.</strong> Do not expect open-weight models to autonomously manage loops, permissions, and tool filtering. Enforce state boundaries deterministically in the host application.</p>
</li>
<li><p><strong>Guard your KV-cache.</strong> Never inject dynamic tool schemas into your primary conversation context. Keep Position 0 static, run extraction in ephemeral hops, and append observation blocks exclusively at the tail.</p>
</li>
</ol>
<hr />
<blockquote>
<h2>📚 Further Reading &amp; Resources</h2>
<p>If you want to dive deeper into the mechanics of local AI agents, constrained decoding, and system architecture, here are the resources that shaped this build:</p>
<ul>
<li><p><a href="https://www.anthropic.com/engineering/building-effective-agents"><strong>Building Effective AI Agents (Anthropic)</strong></a><strong>:</strong> A phenomenal, foundational essay on why developers should avoid complex, multi-agent frameworks and stick to simple, deterministic routing.</p>
</li>
<li><p><a href="https://github.com/ashishpatel26/500-AI-Agents-Projects"><strong>The 500 AI Agents Project</strong></a><strong>:</strong> A massive open-source catalog showcasing how different teams are tackling Agent-Computer Interfaces (ACI).</p>
</li>
<li><p><a href="https://github.com/ollama/ollama"><strong>Ollama Tool Calling Specifications</strong></a><strong>:</strong> The raw documentation on how local models serialize JSON and XML under the hood, which is critical for debugging extraction failures.</p>
</li>
</ul>
<h3>💻 The Code</h3>
<p>Want to run the Guard-Gated Decoupled Dispatch architecture on your own machine? The entire runtime is open-source and runs strictly on local hardware.</p>
<p>Check out the repository, fork it, and let me know how it handles on your GPU:</p>
<p><a href="https://github.com/ShreyasDatta/Local-Autonomous-Agent">https://github.com/ShreyasDatta/Local-Autonomous-Agent</a></p>
</blockquote>
]]></content:encoded></item></channel></rss>