Back to case studies

Case Study

Local AI Coding Agent

A coding agent for local models.

An agentic coding runtime for local, open-weight models, built around inspectable tools, bounded context, and human review.

Private local prototype. Selected architecture and evaluation scope shown.

Independent architecture and TypeScript implementation: agent core, model gateway, context compiler, tool loop, permissions, and evaluation.

Simplified index of system modules
Architecture modules. This does not represent a live run.

Case Snapshot

The strategic brief

The problem, the system response, the available proof, the strategic value, and the intentional boundary.

01Problem
Local coding models need more than a chat interface to inspect a repository, gather evidence, and recover from unproductive tool calls.
02System
An editor-independent TypeScript runtime with a local model gateway, bounded context, repository tools, permission checks, and controlled repair.
03Proof
A documented 24-task read-only evaluation corpus and a separate 30-case deterministic coding baseline. These test different parts of the system.
04Value
Makes local AI-assisted development inspectable at the level of tool choice, evidence, edits, and verification.
05Limitation
A prototype, not a production coding assistant. Deterministic harness results do not establish live model reliability; IDE integration remains planned.

01

Give the model a bounded job

The project separates repository inspection, planning, debugging, and mutation. A model response is only one step: the runtime also needs to know what evidence was gathered, which tools may run, and when the task should stop.

02

A runtime independent of the editor

The TypeScript core connects a context compiler and Markdown project memory to a turn-budgeted agent loop. A gateway normalizes model tool calls; repository tools return bounded observations. Permission checks control edits and verification. A local OpenAI-compatible endpoint supports open-weight inference, with Qwen 2.5 Coder 7B evaluated through Ollama.

03

Test the model and the runtime separately

The read-only corpus contains 24 tasks: 20 development cases and four held-out cases covering repository orientation, symbols, configuration, test diagnosis, and cross-file reasoning. Context configurations of 2,048, 8,192, and 32,768 tokens were investigated. A separate 30-case deterministic baseline exercises edits, verification, and safe rejection. Published summaries disagree on success and tool-selection rates, so no aggregate performance claim is made here.

04

Make failure a visible state

Turn budgets, loop detection, an evidence ledger, and a premature-answer gate constrain non-progress. Mutation uses reviewed patches and base-hash checks; verification and repair remain bounded and subject to human approval. These controls are implemented mechanisms, not a guarantee that every model-driven task succeeds.

05

Local inference, explicit limits

The public case describes selected architecture without exposing the private workspace. The CLI and runtime are implemented; IDE integration is planned. Results from deterministic fixtures are kept separate from live inference. No production adoption, general coding success rate, or comparative latency is claimed.

Related systems

Continue through a related product, proof, or creative-direction logic.