Case Study
Local AI Coding Agent
A coding agent for local models.
An agentic coding runtime for local, open-weight models, built around inspectable tools, bounded context, and human review.
Private local prototype. Selected architecture and evaluation scope shown.
Independent architecture and TypeScript implementation: agent core, model gateway, context compiler, tool loop, permissions, and evaluation.
Case Snapshot
The strategic brief
The problem, the system response, the available proof, the strategic value, and the intentional boundary.
- 01Problem
- Local coding models need more than a chat interface to inspect a repository, gather evidence, and recover from unproductive tool calls.
- 02System
- An editor-independent TypeScript runtime with a local model gateway, bounded context, repository tools, permission checks, and controlled repair.
- 03Proof
- A documented 24-task read-only evaluation corpus and a separate 30-case deterministic coding baseline. These test different parts of the system.
- 04Value
- Makes local AI-assisted development inspectable at the level of tool choice, evidence, edits, and verification.
- 05Limitation
- A prototype, not a production coding assistant. Deterministic harness results do not establish live model reliability; IDE integration remains planned.
01
Give the model a bounded job
The project separates repository inspection, planning, debugging, and mutation. A model response is only one step: the runtime also needs to know what evidence was gathered, which tools may run, and when the task should stop.
02
A runtime independent of the editor
The TypeScript core connects a context compiler and Markdown project memory to a turn-budgeted agent loop. A gateway normalizes model tool calls; repository tools return bounded observations. Permission checks control edits and verification. A local OpenAI-compatible endpoint supports open-weight inference, with Qwen 2.5 Coder 7B evaluated through Ollama.
03
Test the model and the runtime separately
The read-only corpus contains 24 tasks: 20 development cases and four held-out cases covering repository orientation, symbols, configuration, test diagnosis, and cross-file reasoning. Context configurations of 2,048, 8,192, and 32,768 tokens were investigated. A separate 30-case deterministic baseline exercises edits, verification, and safe rejection. Published summaries disagree on success and tool-selection rates, so no aggregate performance claim is made here.
04
Make failure a visible state
Turn budgets, loop detection, an evidence ledger, and a premature-answer gate constrain non-progress. Mutation uses reviewed patches and base-hash checks; verification and repair remain bounded and subject to human approval. These controls are implemented mechanisms, not a guarantee that every model-driven task succeeds.
05
Local inference, explicit limits
The public case describes selected architecture without exposing the private workspace. The CLI and runtime are implemented; IDE integration is planned. Results from deterministic fixtures are kept separate from live inference. No production adoption, general coding success rate, or comparative latency is claimed.