---
title: "Building an Offline AI Coding Agent: Three-Model Architecture, Context OS, and the Plan.MD Pipeline"
description: How ElixirData built a custom AI coding agent using open-source models, a Planner-Generator-Completer pipeline, and Context OS — entirely offline, inside an air-gapped network, with zero data leaving the perimeter.
image: https://www.elixirdata.co/hubfs/elixirdata-og-feature-image.png
---

![campaign-icon](https://assets.elixirdata.co/assets/campaign.svg)

The Context OS for Agentic Intelligence

[![elixir-logo](https://www.elixirdata.co/hubfs/elixirdata-logo.svg)](https://www.elixirdata.co/)

- Platform 
  
    - [Context OS](https://www.elixirdata.co/platform/context-os/)
    - [Unify Data](https://www.elixirdata.co/platform/unify-data/)
    - [Business Context](https://www.elixirdata.co/platform/business-context/)
    - [Decision Infrastructure](https://www.elixirdata.co/platform/decision-infrastructure/)
    - [Build Agents](https://www.elixirdata.co/platform/build-agents/)
    - [Governed Agentic Actions](https://www.elixirdata.co/platform/governed-actions/)
    - [Decision Traces](https://www.elixirdata.co/platform/decisiontraces/)
  
  Platform
  
  The Decision Harness for Enterprise AI.
  
  Three primitives. One dual-gate architecture. Every agent action compiled, governed, and recorded with full lineage.

  [Explore Context OS →](https://www.elixirdata.co/platform/context-os/)
  
  The Three Primitives
  
  [⊞ Context Layer 01 Decision-grade context compiled at the moment of decision](https://www.elixirdata.co/platform/context-os/) [⊛ Governance Layer 02 Dual-gate policy enforcement — before reasoning, before execution](https://www.elixirdata.co/platform/decision-infrastructure/) [◈ Memory Layer 03 Full-lineage Decision Traces, never summarized, never compressed](https://www.elixirdata.co/platform/decisiontraces/) [⟳ Feedback Band Closed-loop improvement across all three layers — 10–17% quarterly accuracy gain](https://www.elixirdata.co/platform/business-context/)

  Agentic Execution
  
  [◉ Build Agents Design and deploy governed agents on Context OS](https://www.elixirdata.co/platform/build-agents/) [▤ Dual-Gate Architecture How every action flows through Gate 1 and Gate 2](https://www.elixirdata.co/platform/governed-actions/) [▦ Decision Traces Live audit record for every agent decision, audit-ready by default](https://www.elixirdata.co/platform/decisiontraces/) [↗ Trust Graduation Shadow → Supervised → Bounded → Full Autonomy](https://www.elixirdata.co/platform/unify-data/)

  **Generic harnesses plateau.** Context OS compounds.
  
  [Download Executive Blueprint →](https://www.elixirdata.co/resources/executive-blueprint/)
- Solutions 
  
    - [Operations & SRE](https://www.elixirdata.co/solutions/operations-sre/)
    - [Security & SOC](https://www.elixirdata.co/solutions/security-and-soc/)
    - [Risk & Compliance](https://www.elixirdata.co/solutions/governance-risk-compliance/)
    - [Finance & Procurement](https://www.elixirdata.co/solutions/finance-and-procurement/)
    - [Agentic Debugging](https://www.elixirdata.co/solutions/agentic-debugging/)
    - [Agentic Code Simulations](https://www.elixirdata.co/solutions/agentic-code-simulations/)
    - [Private AI Assistant with LLM Council](https://www.elixirdata.co/solutions/private-ai-assistant/)
    - [Vision AI and Video Intelligence](https://www.elixirdata.co/solutions/vision-ai/)
    - [Banking & Financial Services](https://www.elixirdata.co/industries/banking-and-financial-services/)
    - [Manufacturing](https://www.elixirdata.co/industries/discrete-manufacturing/)
    - [Transportation](https://www.elixirdata.co/industries/transportation/)
    - [Public Safety](https://www.elixirdata.co/industries/public-safety/)
    - [Travel & Hospitality](https://www.elixirdata.co/industries/travel-and-hospitality/)
    - [Shipping & Logistics](https://www.elixirdata.co/industries/shipping-and-logistics/)
    - [Emergency Services](https://www.elixirdata.co/industries/emergency-services/)
    - [Energy & Utilities](https://www.elixirdata.co/industries/energy-utilities/)
    - [Robotics & Physical AI](https://www.elixirdata.co/industries/robotics-and-physical-ai/)
    - [Industrial Automation](https://www.elixirdata.co/industries/industrial-automation/)
  
  Solutions
  
  Built for regulated enterprise AI.
  
  Find Context OS by the role you own or the industry you operate in. Every solution anchored to the same Decision Harness — governed context, dual-gate enforcement, full-lineage traces.

  [View all solutions →](https://www.elixirdata.co/solutions/operations-sre/)
  
  By Role
  
  [⚖ Risk & Compliance Continuous risk governance with audit-ready Decision Traces](https://www.elixirdata.co/solutions/governance-risk-compliance/) [🛡 Security & SOC Governed threat detection and response with human-in-the-loop authority](https://www.elixirdata.co/solutions/security-and-soc/) [⚡ Operations & SRE Incident response grounded in validated context, every action traced](https://www.elixirdata.co/solutions/operations-sre/) [$ Finance & Procurement Approvals, thresholds, and spend controls enforced before execution](https://www.elixirdata.co/solutions/finance-and-procurement/)

  By Industry
  
  [🏦 Financial Services Model risk management, trading controls, regulatory defensibility](https://www.elixirdata.co/industries/banking-and-financial-services/) [⚕ Healthcare & Life Sciences Clinical decision support, emergency response, life-safety operations](https://www.elixirdata.co/industries/emergency-services/) [🏛 Public Sector Sovereign deployment, tenant isolation, data residency controls](https://www.elixirdata.co/industries/public-safety/) [⚙ Regulated Manufacturing Supply chain intelligence, operational safety, pre-deployment validation](https://www.elixirdata.co/industries/discrete-manufacturing/)

  Don't see your fit? **Every solution is built on the same Decision Harness.**
  
  [Request a custom briefing →](https://www.elixirdata.co/contact-us/)
- Industries 
  
    - [Industries Overview](https://www.elixirdata.co/industries/)
    - [Discrete Manufacturing](https://www.elixirdata.co/industries/discrete-manufacturing/)
    - [Industrial Automation](https://www.elixirdata.co/industries/industrial-automation/)
    - [Robotics & Physical AI](https://www.elixirdata.co/industries/robotics-and-physical-ai/)
    - [Energy & Utilities](https://www.elixirdata.co/industries/energy-utilities/)
    - [Transportation](https://www.elixirdata.co/industries/transportation/)
    - [Shipping & Logistics](https://www.elixirdata.co/industries/shipping-and-logistics/)
    - [Telecommunications](https://www.elixirdata.co/industries/telco/)
    - [Banking & Financial Services](https://www.elixirdata.co/industries/banking-and-financial-services/)
    - [Travel & Hospitality](https://www.elixirdata.co/industries/travel-and-hospitality/)
    - [Public Safety](https://www.elixirdata.co/industries/public-safety/)
    - [Emergency Services](https://www.elixirdata.co/industries/emergency-services/)
  
  Industries
  
  AI Decision Infrastructure for Modern Industry Operations.
  
  Governed, context-aware AI across industrial systems, critical infrastructure, regulated services, and public operations.

  Industrial Systems
  
  [⚙ Discrete Manufacturing Quality, traceability, and production governance](https://www.elixirdata.co/industries/discrete-manufacturing/) [⌘ Industrial Automation Safety boundaries for autonomous industrial systems](https://www.elixirdata.co/industries/industrial-automation/) [◉ Robotics & Physical AI Governed autonomy with human authority](https://www.elixirdata.co/industries/robotics-and-physical-ai/) [⚡ Energy & Utilities Safe, real-time grid decision governance](https://www.elixirdata.co/industries/energy-utilities/)

  Mobility, Networks & Travel
  
  [↗ Transportation Governed transport decisions with full lineage](https://www.elixirdata.co/industries/transportation/) [▦ Shipping & Logistics Routing, asset movement, and traceability](https://www.elixirdata.co/industries/shipping-and-logistics/) [⌁ Telecommunications Accountable AI for network operations](https://www.elixirdata.co/industries/telco/) [✦ Travel & Hospitality Governed, context-aware guest personalization](https://www.elixirdata.co/industries/travel-and-hospitality/)

  Regulated & Public Services
  
  [🏦 Banking & Financial Services Defensible decisions and regulatory controls](https://www.elixirdata.co/industries/banking-and-financial-services/) [🏛 Public Safety Explainable decisions with accountable lineage](https://www.elixirdata.co/industries/public-safety/) [⚕ Emergency Services Governed intelligence for critical response](https://www.elixirdata.co/industries/emergency-services/)

  **Industry-specific operations.** One governed Decision Harness.
  
  [Explore all industries →](https://www.elixirdata.co/industries/)
- Enterprise 
  
    - [Agent Registry](https://www.elixirdata.co/enterprise/agent-registry/)
    - [AgentOps](https://www.elixirdata.co/enterprise/agentops/)
    - [Agent Identity & Access](https://www.elixirdata.co/enterprise/agent-identity-and-access/)
    - [Evaluation and Optimization](https://www.elixirdata.co/enterprise/evaluation-optimization/)
    - [Trust Center](https://www.elixirdata.co/enterprise/trust-center/)
    - [Privacy, Security & Compliance](https://www.elixirdata.co/enterprise/privacy-security-compliance/)
    - [Data Residency & Isolation](https://www.elixirdata.co/enterprise/data-residency/)
    - [Admin & Access Control](https://www.elixirdata.co/enterprise/agent-identity-and-access/)
    - [SLAs & Support](https://www.elixirdata.co/enterprise/ai-sla-support/)
  
  Enterprise
  
  Enterprise control without slowing execution.
  
  Operational governance and compliance-grade trust built into every deployment. Certified to SOC 2, ISO 27001, and defensible under OCC SR 11-7 and the EU AI Act.

  [Visit Trust Center →](https://www.elixirdata.co/enterprise/trust-center/)
  
  Agent Operations
  
  [◉ Agent Registry Approve agents, scopes, tools, and versions with full lifecycle management](https://www.elixirdata.co/enterprise/agent-registry/) [◎ AgentOps Monitor execution, track boundary violations, one-click rollback](https://www.elixirdata.co/enterprise/agentops/) [⚿ Agent Identity Scoped access per task — no over-permissioning, no added risk](https://www.elixirdata.co/enterprise/agent-identity-and-access/) [↗ Trust Graduation Shadow → Supervised → Bounded → Full Autonomy lifecycle](https://www.elixirdata.co/enterprise/evaluation-optimization/)

  Trust & Governance
  
  [⛉ Trust Center SOC 2 · ISO 27001 · CSA STAR · EU AI Act defensibility](https://www.elixirdata.co/enterprise/trust-center/) [⌖ Data Residency & Isolation Region controls, tenant isolation, full data sovereignty](https://www.elixirdata.co/enterprise/data-residency/) [◌ Workforce IAM Roles, SSO, least privilege across humans and AI coworkers](https://www.elixirdata.co/enterprise/agent-identity-and-access/) [◈ SLAs & Support Uptime guarantees, response times, escalation paths](https://www.elixirdata.co/enterprise/ai-sla-support/)

  **Audit-ready by default.** Defensible under regulation.
  
  [Request Trust Package →](https://www.elixirdata.co/enterprise/privacy-security-compliance/)
- Resources 
  
    - [Executive Blueprint](https://www.elixirdata.co/resources/executive-blueprint/)
    - [Blog](https://www.elixirdata.co/blog/)
    - [Customer Outcomes](https://www.elixirdata.co/resources/customer-outcomes/)
    - [Trust & Assurance](https://www.elixirdata.co/trust-and-assurance/authority-model/)
  
  Resources
  
  ### [Executive Blueprint Strategic guide for enterprise AI leaders](https://www.elixirdata.co/resources/executive-blueprint/)
  
  ### [Blog Insights on modern AI systems](https://www.elixirdata.co/blog/)
  
  ### [Customer Outcomes Proof of impact for clients](https://www.elixirdata.co/resources/customer-outcomes/)
  
  ### [Trust and Assurance Framework for governed decisions](https://www.elixirdata.co/trust-and-assurance/)
- Company 
  
    - [About Us](https://www.elixirdata.co/about-us/)
    - [Leadership](https://www.elixirdata.co/leadership/)
    - [Careers](https://www.elixirdata.co/careers/)
    - [Press & News](https://www.elixirdata.co/press-and-news/)
    - [Contact](https://www.elixirdata.co/contact-us/)
    - [Governance and Transparency](https://www.elixirdata.co/governance-and-transparency/)
  
  About
  
  ### [About Us Learn more about our mission and vision](https://www.elixirdata.co/about-us/)
  
  ### [Leadership Meet our experienced executive leadership team](https://www.elixirdata.co/leadership/)
  
  ### [Careers Join us in building enterprise AI solutions](https://www.elixirdata.co/careers/)
  
  ### [Press & News Stay informed with latest company updates](https://www.elixirdata.co/press-and-news/)
  
  ### [Contact Get in touch with our team directly](https://www.elixirdata.co/contact-us/)

  Company
  
  ### Governance and Transparency
  
   Discover the principles, leadership, and culture driving our approach to secure and governed enterprise AI.
  
   Learn how our frameworks for trust, compliance, and operational rigor ensure transparency and accountability at scale. 
  
  [Learn More →](https://www.elixirdata.co/governance-and-transparency)
- [Pricing](https://www.elixirdata.co/pricing/)
  
  [Pricing](https://www.elixirdata.co/pricing/)
  
  ### Pricing Overview
  
  Clear and transparent pricing models
  
  ### Deployment Options
  
  Flexible and scalable cloud choices
  
  ### Enterprise Engagement Model
  
  Customized solutions with tailored pricing

  Our Plans
  
  Pricing Tailored to Your Needs
  
  Learn about our pricing structure, plans, and options tailored to your needs.
  
  View Plans →

[Get Demo](https://www.elixirdata.co/context-os/demo/)

[LLMS TXT](https://www.elixirdata.co/llms.txt) [LLMS Full TXT](https://www.elixirdata.co/llms-full.txt) [AI Context JSON](https://www.elixirdata.co/ai-context.json)

[AI Coding Agent](https://www.elixirdata.co/blog/tag/ai-coding-agent)

# Building an Offline AI Coding Agent: Three-Model Architecture, Context OS, and the Plan.MD Pipeline

[Navdeep Singh Gill](https://www.elixirdata.co/blog/author/navdeep-singh-gill) | 31 March 2026

Building an Offline AI Coding Agent: Three-Model Architecture, Context OS, and the Plan.MD Pipeline

21:34

Three-model architecture, Context OS, and the Plan.MD pipeline

## Key takeaways

- AI coding assistants fail because of context infrastructure, not model capability — 120x token reduction from structured indexing vs file-by-file grep
- A three-model pipeline (Planner → Generator → Completer) decomposes the problem into stages matching how experienced developers actually work
- Plan.MD — a detailed, human-reviewable specification produced before any code is generated — is the highest-leverage artifact in the system
- Context OS orchestrates context across all three tiers: token budget as memory, retrieval as I/O, model routing as process scheduling
- The entire system runs offline in an air-gapped network using open-source models — zero data leaves the perimeter

## Why AI Coding Assistants Fail

Most teams adopting AI coding assistants in 2026 follow the same playbook: subscribe to a cloud API, plug it into the IDE, and hope for the best. The results are often underwhelming. The AI generates code that does not fit the codebase, calls APIs that do not exist, and applies patterns that were deprecated months ago. Developers spend as much time correcting the output as they would have spent writing it themselves.

This cycle is familiar enough that it has a name: the Frustration Loop. Generate code, review it, find it does not fit, regenerate with corrections, review again, eventually accept heavily-modified output or abandon the attempt entirely.

ElixirData faced these problems — and then added another constraint: the entire system has to run offline, inside an air-gapped development network, with zero data leaving the perimeter. No cloud fallback. No external API calls. No live documentation retrieval.

This constraint turned out to be clarifying. It forced the team to confront a truth that cloud-connected teams can afford to ignore: the quality of an AI coding assistant is determined not by the model it runs, but by the context infrastructure underneath it — and more specifically, by the quality of the plan the model works from.

**The thesis: The model is the engine. Context is the fuel. The plan is the route. Better fuel and a better route produce better output — regardless of the engine.**

### Two context gaps, same failure mode

A typical AI coding assistant operates with two sources of knowledge: training data (static, frozen at a cutoff date) and whatever context fits in the current prompt (dynamic but shallow). This creates two distinct failures that produce the same downstream result: bad code.

**Internal context failure:** The assistant does not understand the codebase. It does not know the project uses Fastify instead of Express, that files go in lib/services/ instead of utils/, or that the codebase follows a functional style. When it needs to understand the authentication flow, it reads files one by one, burning hundreds of thousands of tokens for what a structured index could answer in a few hundred.

**External context failure:** The assistant does not have current knowledge of libraries and APIs. It falls back on training data — which may be months or years stale. For fast-moving libraries, this means code that calls methods that no longer exist, passes parameters in the wrong order, or misses simpler approaches added in recent versions.

Project-reported benchmarking from codebase-memory-mcp shows approximately 3,400 tokens via structured indexing versus approximately 412,000 tokens via file-by-file grep for the same five structural queries — roughly a 120x reduction.

### The cascade problem

These failures compound. As Dex Horthy of HumanLayer articulated: a bad line of code is a bad line of code, but a bad line in a plan leads to hundreds of bad lines of code, and bad research leads to thousands. Upstream errors cascade and amplify downstream.

If planning quality is the primary lever — if the accuracy of the plan determines the quality of everything downstream — then the architecture should optimize for planning quality above all else.

### The speed trap

Part of the problem lies in how success is measured. Teams often track time to first output or lines of code generated. If an AI generates 200 lines in seconds but a developer spends 30 minutes refactoring them, the net productivity gain is questionable.

| Misleading metric | More useful alternative |
| --- | --- |
| Time to first output | First-pass acceptance rate |
| Lines of code generated | Iteration cycles per task |
| Tasks completed | Post-merge rework required |
| Generation speed | Review burden vs manual writing |

## The Three-Model Architecture

ElixirData's solution inverts the conventional single-model approach. Three models operate in a pipeline, each optimized for a different cognitive task.

`Request → [Planner] → Plan.MD → [Generator] → Code → [Completer] → Edits`

### Tier 1: The planner

| Property | Specification |
| --- | --- |
| Model | Qwen3-Coder-480B (35B active parameters, MoE) |
| Role | Research codebase, reason about architecture, produce Plan.MD |
| License | Apache 2.0 |
| Context window | 256K–1M tokens |
| Latency target | 30–60 seconds (thoroughness over speed) |
| Token budget | Up to 128K (deep codebase understanding) |
| Cloud alternative | Claude Opus 4.6 for non-sensitive work outside the air gap |

The planner is the most important model in the system. Its job is to research the codebase, understand the request, reason about architectural implications, and produce a detailed Plan.MD that specifies exactly what to build, where to build it, which patterns to follow, and which edge cases to handle.

For teams with cloud access for non-sensitive work, Claude Opus 4.6 is the ideal planner. Its adaptive thinking, 500K–1M token context window, and frontier-level reasoning make it the strongest option available. Use it outside the air gap for architecture research and specification generation, then transfer the outputs via secure media.

### Tier 2: The generator

| Property | Specification |
| --- | --- |
| Model | GLM-5 (744B total, 40B active parameters, MoE) |
| Role | Execute Plan.MD — write the actual code |
| License | MIT |
| SWE-bench Verified | 77.8% (best open-source) |
| Latency target | \< 15 seconds first token |
| Token budget | Up to 64K (Plan.MD has narrowed the scope) |
| Alternatives | Kimi K2.5 (99% HumanEval, 32B active) or DeepSeek-V3.2 (MIT) |

The generator receives Plan.MD and executes against it. It does not need to reason about architecture — the planner already did that. It translates a precise specification into correct, convention-following code.

When Plan.MD is accurate — when it correctly identifies files, patterns, edge cases, and conventions — even a mid-tier generation model will produce correct code. The plan has already done the hard cognitive work.

### Tier 3: The completer

| Property | Specification |
| --- | --- |
| Model | Qwen3-Coder-Next (80B total, 3B active parameters, MoE) |
| Role | Fast inline completions and small edits |
| License | Apache 2.0 |
| SWE-bench | 70.6% (comparable to models 10–20x larger) |
| Latency target | \< 500 milliseconds |
| Token budget | Up to 8K (immediate editing context only) |
| What it handles | Ghost text, autocomplete, FIM, single-function refactors |

The completer handles 80% of daily interactions by volume. It does not need Plan.MD or full context retrieval — it needs the current file, the imports, and the path-scoped rules.

### Plan.MD: The contract between stages

Plan.MD is the critical artifact in the pipeline. It serves as an explicit, inspectable, human-reviewable contract between the planner and the generator. A well-formed Plan.MD includes:

**Research findings:** What the planner learned about the codebase (with file paths and line numbers)

**Files to create and modify:** Specific locations, not vague module references

**Architecture decisions:** Chosen approach with rationale

**Edge cases to handle:** Failure modes, boundary conditions, exceptional flows

**Test plan:** What to test and how

**Conventions to follow:** Referenced from CLAUDE.md with specific examples

The specificity matters. "Modify the auth module" is not a plan. "Add rate limiting check in src/gateway/middleware.py after line 47, following the decorator pattern used in auth.py:23-47" is a plan. The more specific the plan, the less the generator needs to guess — and guessing is where models fail.

Critically, Plan.MD is a human-readable artifact. A developer can review it before generation begins and catch errors at the plan stage, where they are cheap to fix, rather than at the code stage, where they cascade.

## The Context Operating System

The Context OS is the orchestration layer that ensures each model sees the right context at the right time. It manages three resources that behave like operating system primitives.

### Three resources, three management strategies

**Token budget as memory:** A model's context window is finite. The Context OS allocates space to the most relevant context, evicts stale information, and ensures totals stay within limits. Dumping everything degrades performance the same way memory swapping does.

**Retrieval as I/O:** Fetching context from the codebase index, documentation packs, and convention rules is analogous to disk I/O. The Context OS caches frequently-accessed context, prefetches likely-needed documentation, and batches retrieval calls.

**Model routing as process scheduling:** Completion requests need speed (sub-500ms). Generation requests need quality (15 seconds). Planning requests need thoroughness (30+ seconds). The Context OS routes each request to the appropriate model.

### The context pipeline

When a developer makes a request, the Context OS executes six steps:

1. **Request classification:** Determine request type. Completions go directly to Tier 3. Generation and review enter the full pipeline.
2. **Convention loading:** Load CLAUDE.md (always on), path-scoped rules for relevant file types, and matching skills.
3. **Codebase retrieval:** Two-stage: lexical search (ripgrep, milliseconds, 60–70% of queries) then semantic search (Qdrant vectors, 30–40%).
4. **Documentation retrieval:** Pull version-pinned library documentation for dependencies involved.
5. **Context assembly:** Merge, rank by relevance, trim to token budget. Priority: conventions \> code spans \> docs \> broader context.
6. **Enriched inference:** Assembled context combined with request and sent to the appropriate model.

### The three context subsystems

**Codebase index:** Persistent representation of the entire codebase using Tree-sitter for AST parsing and Qdrant for vector storage. Code split at meaningful boundaries — functions, classes, logical blocks. Updated incrementally via Git hooks on every merge to main.

**Documentation packs:** Curated, version-pinned documentation for key dependencies as markdown files optimized for LLM consumption. Also includes actual library source code for critical dependencies. Version-pinned to exact lockfile versions.

**Convention rules:** Structured encoding of coding standards following Claude Code's hierarchy: CLAUDE.md (always loaded), path-scoped rules (loaded per file type), and skills (lazy-loaded bundles of instructions and resources).

## Building from Claude Code's System Prompts

Claude Code's system prompts are fully documented in the open-source repository Piebald-AI/claude-code-system-prompts. As of March 2026 (version 2.1.84), this includes the core system prompt (2,896 tokens), 18 built-in tool descriptions, plan subagent prompt (636 tokens), explore subagent prompt (494 tokens), task subagent prompt (294 tokens), approximately 40 system reminders, and tracking across 133 versions.

### Mapping Claude Code subagents to the pipeline

| Claude Code component | ElixirData mapping | Implementation |
| --- | --- | --- |
| Plan subagent (636 tok) | Tier 1 — Planner | Adapted as system prompt for Qwen3-Coder-480B. Produces Plan.MD as standalone artifact. |
| Explore subagent (494 tok) | Tier 1 — Planner explore mode | Read-only research mode for understanding codebase structure before planning. |
| Task subagent (294 tok) | Tier 2 — Generator | Receives Plan.MD sections as individual tasks. Executes with TodoWrite tracking. |
| Core prompt (2,896 tok) | All tiers — adapted per tier | Conciseness mandate, tool-first approach, file reference format across all models. |

### Key patterns adapted from Claude Code

**Conciseness mandate:** Claude Code instructs the model to "answer concisely with fewer than 4 lines unless the user asks for detail" and to "minimize output tokens." Directly adopted — verbose output wastes developer time and context window space.

**TodoWrite pattern:** Claude Code forces the agent to maintain a task list. ElixirData implements this at the pipeline level: Plan.MD is converted into a todo list that the generator works through sequentially, with real-time progress visible to the developer.

**Hooks system:** Claude Code triggers scripts at lifecycle events. ElixirData implements: post-edit hooks (run linter), post-generation hooks (trigger tests), post-session hooks (log to feedback flywheel), and index-update hooks (re-index after merge).

The meta-strategy: Use Claude itself (via API, outside the air gap for non-sensitive work) to write and refine the CLAUDE.md, rules, skills, and system prompts that the local models execute. Using the strongest model to configure the others is the highest-leverage approach.

### The convention hierarchy

**CLAUDE.md (always loaded):** Project-wide conventions — tech stack with versions, directory structure, naming conventions, coding standards. Start with 5–10 conventions and build iteratively.

**Path-scoped rules (loaded per file type):** Convention guidance triggered by file glob patterns. Example: \*.test.ts triggers test-specific rules about Vitest. Keeps always-loaded CLAUDE.md small.

**Skills (lazy-loaded on demand):** Rich bundles of instructions, documentation, scripts, and examples. Each has a SKILL.md with a description the model uses to decide when to load it.

## Infrastructure

### Hardware specification

| Component | Specification | Purpose |
| --- | --- | --- |
| GPU nodes (Planner + Gen) | 2x NVIDIA A100 80GB or H100 | One for planner, one for generator. MoE architectures use VRAM efficiently. |
| GPU node (Completer) | 1x A100 80GB or RTX 4090 24GB | Dedicated to inline completions. RTX 4090 is the cost-optimized option. |
| CPU server | 32+ cores, 128GB RAM, 2TB NVMe | Context OS: indexer, Qdrant, API gateway, monitoring, embedding model. |
| Storage | NAS, 5TB minimum | Model weights, index persistence, audit logs. |
| Network | 10GbE between nodes | Low-latency internal communication. |

Total hardware estimate: $50,000–$120,000 depending on GPU selection. The software stack is entirely open-source: vLLM for inference, Tree-sitter + Qdrant for indexing, ripgrep for search, OpenTelemetry for monitoring.

### Security architecture

**Network isolation:** Dedicated VLAN, no route to internet. All processing stays on-premise.

**Authentication:** LDAP/Active Directory integration. API tokens scoped to team and repository access.

**Audit logging:** All prompts and responses logged to tamper-evident store.

**Model provenance:** Cryptographic hash verification at deployment. Chain of custody for all weights.

**Code isolation:** No execution outside inference sandbox. Tool use mediated with explicit permissions.

## The Feedback Flywheel

The system includes a closed-loop improvement mechanism that compounds quality over time, even without model upgrades.

**Capture:** Every interaction is logged — the request, Plan.MD, generated code, and the developer's response (accept / modify / reject).

**Analyze:** Weekly automated analysis identifies recurring failure patterns. Example: "73% of rejections in payments module involve incorrect webhook handling."

**Improve:** Failure patterns become targeted context improvements: planner failures → better CLAUDE.md or skills; generator failures → specific path-scoped rules; library misuse → updated documentation packs.

**Measure:** Track first-pass acceptance rate by tier. Planner acceptance dropping? Planner context needs attention. Generation dropping with good plans? Generator rules need tuning.

**Why it compounds:** every convention rule added is a failure mode eliminated permanently. Every documentation pack updated is a class of hallucination removed. After 90 days, the rules file will be smaller, more targeted, and more effective than anything written on day one.

## Nine Best Practices

1. **Invest in planning quality above all else.** If Plan.MD is accurate, the choice of generation model barely matters. A detailed, correct plan with specific file references, pattern examples, and edge case coverage will produce good code from any competent model.
2. **Make Plan.MD a reviewable artifact.** Plan.MD should be visible to the developer before generation begins. Correcting a wrong plan takes seconds; correcting wrong code from a wrong plan takes minutes or hours.
3. **Use a three-model strategy.** Completion, generation, and planning are fundamentally different cognitive tasks. Running all three on a single model forces compromises in every direction. MoE architectures make multi-model strategies practical.
4. **Build Claude Code patterns into local models.** Claude Code's system prompts represent thousands of hours of prompt engineering. Do not reinvent them — adapt them. Use Claude itself to write the prompts your local models execute.
5. **Build the rules file iteratively.** Start with 5–10 conventions. Deploy. Observe rejections. Add targeted rules. Remove rules that never trigger. After 90 days the file will be organically tuned to actual failure modes.
6. **Version-pin everything.** Every documentation pack must be version-pinned to the exact library version in the lockfile. Stale documentation — even by one minor version — can be worse than none.
7. **Measure first-pass acceptance rate.** Lines generated and tokens processed are operational metrics. First-pass acceptance rate captures whether the system produces useful code. Target: above 60% after 6 months.
8. **Deploy in three phases.** Phase 1 (weeks 1–4): Completer only with 5–8 pilot developers. Phase 2 (weeks 5–10): Full pipeline with Plan.MD. Phase 3 (weeks 11–16): Feedback flywheel operational, skills library, performance tuning.
9. **Plan for quarterly model refreshes.** The open-source landscape evolves rapidly. Each refresh: validate on a subset, compare acceptance rates, cut over only when the new model matches or exceeds production performance.

## Conclusion

The most important property of this architecture is that it compounds over time, even without model upgrades. Every convention rule added eliminates a failure mode permanently. Every Plan.MD template refined produces better plans. Every feedback cycle converts developer frustration into systematic improvement.

The open-source ecosystem has delivered models capable of running this pipeline today. Claude Code's published system prompts provide the behavioral blueprint. The three-model architecture decomposes the problem into stages matching how experienced developers actually work. And Plan.MD — a simple markdown file specifying what to build before anything gets built — turns out to be the highest-leverage artifact in the entire system.

**The thesis: The model is the engine. Context is the fuel. The plan is the route. Better fuel and a better route produce better output — regardless of the engine.**

## References

- Claude Code System Prompts: github.com/Piebald-AI/claude-code-system-prompts (v2.1.84, March 2026)
- Qwen3-Coder: github.com/QwenLM/Qwen3-Coder (Apache 2.0, 480B / 35B active)
- GLM-5: Zhipu AI (MIT License, 744B / 40B active, 77.8% SWE-bench)
- Qwen3-Coder-Next: Alibaba (Apache 2.0, 80B / 3B active)
- Kimi K2.5: Moonshot AI (MIT License, 1T / 32B active, 99% HumanEval)
- DeepSeek-V3.2: DeepSeek (MIT License, 685B / 37B active)
- Advanced Context Engineering: Dex Horthy, HumanLayer, August 2025
- Codified Context: Aristidis Vasilopoulos, arXiv, February 2026
- Coding Agents in Feb 2026: Calvin French-Owen, February 2026

## Related Resources

- Defining Context OS — The Definitive Article
- Context Engineering Is Necessary But Not Sufficient
- Your AI Agent Is Failing in Production: 9 Reasons, None Are the LLM
- What Is Context OS? — The Complete Guide
- The Decision Gap: Why Enterprise AI Agents Fail in Production

## See Context OS in action

Request an Executive Briefing to see how Context OS governs AI agents from coding to enterprise execution.

[Request Executive Briefing → demo.elixirdata.co](https://demo.elixirdata.co)

*ElixirData | Context OS™ | The governed operating system for enterprise AI agents*

## Share Article

- [![XenonStack Facebook](https://www.xenonstack.com/hubfs/xenonstack-facebook-service.svg)](http://www.facebook.com/share.php?u=https://www.elixirdata.co/blog/building-an-offline-ai-coding-agent)
- [![XenonStack Twitter](https://www.xenonstack.com/hubfs/xs-twitter-white-updated-icon.svg)](https://twitter.com/intent/tweet?text=I+found+this+interesting+blog+post&url=https://www.elixirdata.co/blog/building-an-offline-ai-coding-agent)
- [![XenonStack Linked In](https://www.xenonstack.com/hubfs/xenonstack-linkedin-service.svg)](http://www.linkedin.com/shareArticle?mini=true&url=https://www.elixirdata.co/blog/building-an-offline-ai-coding-agent)
- [![XenonStack Email Icon](https://www.xenonstack.com/hubfs/xenonstack-email-service.svg)](mailto:?subject=Check%20out%20https://www.elixirdata.co/blog/building-an-offline-ai-coding-agent%20&body=Check%20out%20https://www.elixirdata.co/blog/building-an-offline-ai-coding-agent)

## Table of Contents

## Explore Related Topics

[Agentic Operations](https://www.elixirdata.co/blog/tag/agentic-operations)

[Context Application](https://www.elixirdata.co/blog/tag/context-application)

[Context Graph](https://www.elixirdata.co/blog/tag/context-graph)

[Context OS](https://www.elixirdata.co/blog/tag/context-os)

[Decision Graph](https://www.elixirdata.co/blog/tag/decision-graph)

[Knowledge Graph](https://www.elixirdata.co/blog/tag/knowledge-graph)

[Ontology](https://www.elixirdata.co/blog/tag/ontology)

![navdeep-singh-gill](https://www.elixirdata.co/hubfs/Imported%20images/navdeep-gill-ceo-xenonstack.svg)

## Navdeep Singh Gill

Global CEO and Founder of ElixirData

Navdeep Singh Gill is serving as Chief Executive Officer and Product Architect at XenonStack. He holds expertise in building SaaS Platform for Decentralised Big Data management and Governance, AI Marketplace for Operationalising and Scaling. His incredible experience in AI Technologies and Big Data Engineering thrills him to write about different use cases and its approach to solutions.

[Explore More by Navdeep Singh Gill ![cta-blue-arrow](https://www.elixirdata.co/hubfs/Imported%20images/cta-arrow-blue.svg)](https://www.elixirdata.co/blog/author/navdeep-singh-gill)

## Subscribe to our Latest Technology Insights and Resources

Subscribe Now

![slider-cross-icon](https://www.xenonstack.com/hubfs/slider-cross-icon.svg)

## Get the latest articles in your inbox

Business Email ID \*

Please enter a valid Business Email ID

Company Name \*

Please enter a valid Company Name

Yes, I would like to receive the ElixirData newsletter as well as marketing emails regarding ElixirData products, services, and events. I understand I can unsubscribe at any time.   
By registering, I confirm that I agree to the processing of my personal data by ElixirData as described in the Privacy Policy.

Subscribe Now

## Related Articles for you

![Building an Offline AI Coding Agent: Three-Model Architecture, Context OS, and the Plan.MD Pipeline](https://www.elixirdata.co/hubfs/Xenon%20Daily%20Work-1%20(82).png)

### [Building an Offline AI Coding Agent: Three-Model Architecture, Context OS, and the Plan.MD Pipeline](https://www.elixirdata.co/blog/building-an-offline-ai-coding-agent)

31 March 2026

![Semantic Layer for AI Agents: Why Dashboards Aren't Enough](https://www.elixirdata.co/hubfs/Xenon%20Daily%20Work-1%20%2898%29.png)

### [Semantic Layer for AI Agents: Why Dashboards Aren't Enough](https://www.elixirdata.co/blog/semantic-layer-ai-agents-decision-grade)

22 September 2026

![AI Agent Reliability Enterprise | Stop Duplicate Actions](https://www.elixirdata.co/hubfs/Xenon%20Daily%20Work-1%20-%202026-04-08T131006.122.png)

### [AI Agent Reliability Enterprise | Stop Duplicate Actions](https://www.elixirdata.co/blog/ai-agent-reliability-enterprise-idempotency)

08 April 2026

![elixir-logo](https://www.elixirdata.co/hubfs/elixirdata-logo.svg)

ElixrData is the Decision Harness for Enterprise AI agents. Context tells AI what's true. Governance tells AI what's allowed.

[Get Demo](https://www.elixirdata.co/context-os/demo/)

### Platform

[Context OS](https://www.elixirdata.co/platform/context-os/) [Build Agents](https://www.elixirdata.co/platform/build-agents/) [Unify Data](https://www.elixirdata.co/platform/unify-data/) [Business Context](https://www.elixirdata.co/platform/business-context/) [Decision Infrastructure](https://www.elixirdata.co/platform/decision-infrastructure/) [Agentic Actions](https://www.elixirdata.co/platform/governed-actions/) [Decision Traces](https://www.elixirdata.co/platform/decisiontraces/)

### Solutions

[Operations & SRE](https://www.elixirdata.co/solutions/operations-sre/) [Security & SOC](https://www.elixirdata.co/solutions/security-and-soc/) [Risk & Compliance](https://www.elixirdata.co/solutions/governance-risk-compliance/) [Finance & Procurement](https://www.elixirdata.co/solutions/finance-and-procurement/) [Agentic Debugging](https://www.elixirdata.co/solutions/agentic-debugging/) [Vision AI](https://www.elixirdata.co/solutions/vision-ai/)

All industries

### Enterprise

[Agent Registry](https://www.elixirdata.co/enterprise/agent-registry/) [AgentOps](https://www.elixirdata.co/enterprise/agentops/) [Agent Identity & Access](https://www.elixirdata.co/enterprise/agent-identity-and-access/) [Evaluation & Optimization](https://www.elixirdata.co/enterprise/evaluation-optimization/) [Trust Center](https://www.elixirdata.co/enterprise/trust-center/) [Data Residency](https://www.elixirdata.co/enterprise/data-residency/) [SLAs & Support](https://www.elixirdata.co/enterprise/ai-sla-support/)

### Integrations

[Databricks](https://www.elixirdata.co/integrations/databricks/) [Looker](https://www.elixirdata.co/integrations/looker/) [Power BI](https://www.elixirdata.co/integrations/power-bi/) [Qlik](https://www.elixirdata.co/integrations/qlik/) [AWS QuickSight](https://www.elixirdata.co/integrations/aws-quicksight/) [SAP](https://www.elixirdata.co/integrations/sap/) [Sigma Computing](https://www.elixirdata.co/integrations/sigma-computing/) [Snowflake](https://www.elixirdata.co/integrations/snowflake/) [Spotfire](https://www.elixirdata.co/integrations/spotfire/) [Tableau](https://www.elixirdata.co/integrations/tableau/) [ThoughtSpot](https://www.elixirdata.co/integrations/thoughtspot/) [Traditional Analytics](https://www.elixirdata.co/integrations/traditional-analytics/)

### Resources

[Executive Blueprint](https://www.elixirdata.co/resources/executive-blueprint/) [Blog](https://www.elixirdata.co/blog/) [Customer Outcomes](https://www.elixirdata.co/resources/customer-outcomes/) [Trust and Assurance](https://www.elixirdata.co/trust-and-assurance/)

### Company

[About Us](https://www.elixirdata.co/about-us/) [Leadership](https://www.elixirdata.co/leadership/) [Careers](https://www.elixirdata.co/careers/) [Press & News](https://www.elixirdata.co/press-and-news/) [Contact](https://www.elixirdata.co/contact-us/)

© 2026 ElixirData | Context OS™ — Making Context Executable, Enforceable, and Governed

Privacy

Terms

Security

Cookies

[LLMS TXT](https://www.elixirdata.co/llms.txt) [LLMS Full TXT](https://www.elixirdata.co/llms-full.txt) [AI Context JSON](https://www.elixirdata.co/ai-context.json)

[Telco](https://www.elixirdata.co/industries/telco/) [Agent Ecosystem](https://www.elixirdata.co/ai-agents/agent-ecosystem/) [Hyperautomation Generative AI Book](https://www.elixirdata.co/newsroom/press-release/hyperautomation-generative-ai-book/) [Resources](https://www.elixirdata.co/resources/) [Audit Agent](https://www.elixirdata.co/ai-agents/audit-agent/) [Approval Agent](https://www.elixirdata.co/ai-agents/approval-agent/) [Decision Review Agent](https://www.elixirdata.co/ai-agents/decision-review-agent/) [Enterprise](https://www.elixirdata.co/enterprise/) [Compliance Agent](https://www.elixirdata.co/ai-agents/compliance-agent/) [Exception Handling Agent](https://www.elixirdata.co/ai-agents/exception-handling-agent/) [ElixirOS](https://www.elixirdata.co/product/elixiros/)

```json
{
  "@context" : "https://schema.org",
  "@type" : "BlogPosting",
  "author" : {
    "@type" : "Person",
    "name" : "Navdeep Singh Gill",
    "url" : "https://www.elixirdata.co/blog/author/navdeep-singh-gill"
  },
  "dateModified" : "2026-03-31T06:23:22.453Z",
  "datePublished" : "2026-03-30T04:17:05.000Z",
  "headline" : "Building an Offline AI Coding Agent: Three-Model Architecture, Context OS, and the Plan.MD Pipeline",
  "image" : [ "https://www.elixirdata.co/hubfs/Xenon%20Daily%20Work-1%20(82).png" ],
  "mainEntityOfPage" : {
    "@id" : "https://www.elixirdata.co/blog/building-an-offline-ai-coding-agent",
    "@type" : "WebPage"
  },
  "publisher" : {
    "@type" : "Organization",
    "logo" : {
      "@type" : "ImageObject"
    },
    "name" : "ElixirData"
  }
}
```