All notes

#209 — Harness engineering

September 26, 2026·4 min read

#209 — Harness engineering

AI coding agents can generate code faster than human teams, but without structural constraints they swamp your senior engineers in review fatigue and technical debt. Unlocking true leverage requires treating code generation as a control systems problem: your engineering team must build an outer "harness" that steers agent behavior up front and forces automated self-correction before code ever reaches human eyes.

The Core Equation

A production coding agent is not just a language model; it is the model wrapped in an execution environment.

  • The formula: Agent = Model + Harness
  • Inner vs. outer harness: The model vendor provides the inner harness (e.g., base system instructions, context window management, and native tool execution). Your team must engineer the outer harness (e.g., repository rules, architecture boundaries, custom linters, and verification loops).
  • The root challenge: LLMs operate probabilistically over tokens without semantic comprehension, social accountability, or institutional memory. Trust requires externalizing implicit developer knowledge into explicit systemic constraints.

The Cybernetic Loop: Guides vs. Sensors

Harness engineering borrows from control theory, operating as a continuous cybernetic governor. Every engineering intervention falls into one of two control mechanisms:

  • Guides (feedforward controls): Steer the agent before it generates code, anticipating failure modes to narrow the model's solution space and ensure first-pass accuracy.
  • Sensors (feedback controls): Observe artifacts after generation, feeding actionable error diagnostics directly back into the agent so it can self-heal without human intervention.

Control Matrix: Computational vs. Inferential

Every guide and sensor executes across two distinct computational layers. Computational tools provide cheap, deterministic guardrails, while inferential tools apply flexible, semantic judgment.

Harness DimensionComputational Controls (Deterministic, fast, runs on CPU)Inferential Controls (Probabilistic, semantic, runs on GPU)
Feedforward (Guides)Language Server Protocol (LSP), AST scaffolding, starter templates, deterministic CLI toolsAGENTS.md / CLAUDE.md context files, curated few-shot examples, domain guidelines, design system heuristics
Feedback (Sensors)Fast compilers, type checkers, deterministic linters, unit tests, structural architecture tests (e.g., ArchUnit)Secondary LLM review agents, automated code-judge prompts, pull request risk evaluators

The Three Quality Tiers

You cannot leash an agent with a single generic test suite. A complete harness manages three distinct quality dimensions:

  • Maintainability: Governs style, cyclomatic complexity, deduplication, and code aesthetics. This tier is almost entirely solved via fast, deterministic computational tools like linters and formatters.
  • Architecture fitness: Enforces modular boundaries, layer isolation, data contracts, and dependency rules. It requires codifying fitness functions (custom lint rules, dependency drift detectors) so agents cannot break system invariants.
  • Functional behavior: Verifies that the software actually satisfies business requirements. This is the hardest tier to automate; relying on an agent to both write business logic and generate its own test suite creates an unverified, circular feedback loop.

Where Sensors Run

Feedback loops must operate across multiple operational cadences to balance latency, cost, and safety:

  • In-session loops: Sub-second type-checks, linters, and unit test suites running in the agent's local terminal.
  • Continuous integration: End-to-end integration runs, contract checks, and LLM-as-judge diff reviews triggered prior to merge.
  • Scheduled background sweeps: Asynchronous "garbage collector" agents that continuously hunt for stale context, documentation drift, dead dependencies, and architectural erosion across the repository.
  • Production telemetry: Runtime observability, synthetic probes, and automated error tracking looped back as context into the development environment.

The Founder Playbook

  • Stop prompting; start building harnesses: Treat agent failures, hallucinated libraries, or broken patterns as infrastructure deficiencies. When an agent makes a mistake, codify the fix into a machine-readable rule file or deterministic test.
  • Select opinionated, strongly typed stacks: Dynamic, unconstrained languages dramatically increase the model's error rate. Strongly typed languages paired with strict compiler feedback provide an instant, zero-cost computational guide rail.
  • Build harness templates for core topologies: Identify repetitive architectural patterns across your product (e.g., event consumers, CRUD services, data pipelines). Package each pattern into a reusable template pre-fitted with boilerplate guides, schema contracts, and self-testing suites.
  • Never deploy a sensor without an automated closed loop: A sensor that simply alerts a human creates review fatigue. Pipe structured diagnostic output directly back into the agent's context window so it autonomously iterates until tests pass.
  • Retain human judgment at the boundary: Use automation to absorb maintainability and architectural checks, but keep your senior engineers focused strictly on behavioral intent, threat modeling, and product edge cases.

The Bottom Line

Unconstrained coding agents do not reduce engineering headcount or timeline; they convert writing time into review time. Real leverage requires building deterministic compilers, fixed API contracts, and automated test loops that force models to self-correct before code hits human review.

If any of this resonates with what you’re building, or if you’re considering working together, you can reach us here.

Frequently asked questions

What is harness engineering, and why does my startup need it if AI models keep getting smarter?

Harness engineering is the practice of designing feedforward constraints (guides) and feedback loops (sensors) around coding agents so they produce reliable, production-ready code. Smarter LLMs increase generation speed, but they do not possess institutional context, architectural intuition, or accountability. As documented by teams at OpenAI and Stripe, raw intelligence shifts your engineering bottleneck from writing code to human review toil unless the execution environment automatically enforces codebase conventions and self-correction.

How does harness engineering differ from standard prompt engineering or context engineering?

Prompt engineering optimizes instructions inside a prompt window, and context engineering retrieves relevant tokens for that window. Harness engineering builds the persistent, cybernetic operating environment across both CPU and GPU tooling. It combines deterministic static tools (linters, AST codemods, language server protocols) with semantic evaluators to govern code before generation and trigger automated self-healing loops after execution.

What is the difference between computational and inferential controls in an agent harness?

Computational controls are deterministic, execute in milliseconds on a CPU, and cost virtually nothing (e.g., TypeScript compilers, ArchUnit boundary tests, ESLint). Inferential controls are probabilistic, run on GPUs or NPUs, and handle subjective nuance (e.g., LLM-as-a-judge reviews, semantic diff inspections). Early-stage engineering teams prioritize computational controls because they provide cheap, immediate failure signals that agents can reliably parse and fix in-session.

Why do AI coding agents struggle with functional behavior testing, and how do we solve the circular testing trap?

Allowing an agent to write both application logic and the corresponding test suite creates a circular validation loop where the AI writes flawed tests to validate its own flawed assumptions. Engineering teams mitigate this by separating the spec writer from the test author, introducing approved fixtures (curated snapshots of known good inputs and outputs), and using mutation testing (like Stryker or PIT) to verify whether test suites detect deliberately injected bugs.

How should technical founders evaluate harnessability when selecting an engineering tech stack?

Harnessability is the degree to which an environment provides ambient affordances (legibility and tractability) to an autonomous agent. Stacks with strong static typing (such as TypeScript, Go, or Rust), explicit module boundaries, and mature scaffolding tooling provide immediate feedback channels for agent self-correction. Dynamically typed, unconstrained codebases lack built-in computational rails, forcing you to rely on expensive and probabilistic inferential judges.

What are harness templates, and how do they reduce architectural drift in scaling teams?

Harness templates are pre-packaged bundles of scaffolding, architectural fitness functions, and feedback sensors tailored to common system topologies like CRUD APIs or asynchronous event workers. Under Ashby's Law of Requisite Variety, an unconstrained agent will explore infinite ways to build a service. Templates restrict the solution space to a vetted pattern, allowing startups to scale microservices or features without accumulating inconsistent boilerplate.

How did companies like Stripe and OpenAI implement autonomous agent harnesses in production?

OpenAI built layered architectures validated by custom linters and structural tests, backed by automated background 'janitor' agents that scan repositories for architectural drift and open remediation PRs. Stripe implemented Minions, an end-to-end coding agent framework using blueprint-driven feedforward constraints and predictive pre-push linter heuristics that feed diagnostics back to the agent before code reaches human code review.

What are ambient affordances in software architecture, and how do they impact coding agents?

Coined by Ned Letcher, ambient affordances describe structural properties of an environment that make code legible, navigable, and tractable for an autonomous agent. Monolithic frameworks with opinionated directory conventions (like Ruby on Rails or Django) and explicit language server protocols provide strong ambient affordances out of the box, whereas bespoke, loosely coupled microservices force teams to build heavy custom tooling just to make the repository interpretable.

How does Ashby's Law of Requisite Variety explain why unconstrained coding agents fail?

Ashby's Law states that a control system can only regulate what it has a model of, and its regulator must possess at least as much variety as the disturbances it governs. Because modern LLMs can generate an infinite variety of code patterns, an unconstrained agent inevitably escapes quality control. Harness engineering acts as a variety-reduction mechanism, using rigid templates, API schemas, and architectural boundaries to keep the agent's problem space strictly within the regulator's capacity.

What role should human software engineers play in an agent-harnessed startup?

Instead of line-by-line syntax reviews, human engineers transition into control systems designers and architectural arbiters. Senior developers externalize their implicit taste, security assumptions, and domain constraints into explicit guides and sensors, while reserving manual review exclusively for high-stakes business logic, product trade-offs, and behavioral intent that automated tools cannot inspect.

Why are mutation testing and structural architecture testing seeing a resurgence alongside coding agents?

Traditional unit test coverage metrics are easily gamed by LLMs that write low-value assertions simply to make coverage turn green. Mutation testing (which injects intentional bugs to see if tests fail) and structural tests (like ArchUnit or Dependency-Cruiser) provide objective computational verification that agent-authored tests actually catch regressions and strictly adhere to layer isolation boundaries.

How do you prevent harness drift as custom rules, sensors, and prompts multiply across a repository?

Harness drift occurs when competing prompt instructions, stale configuration files, and outdated linter rules give contradictory guidance to an agent. High-performing engineering teams treat the harness itself as code: rules files (AGENTS.md) and sensor suites are versioned in Git, tested continuously against synthetic benchmark tasks, and pruned periodically using automated cleanup sweeps.

more than just words|

If you’re considering working together, you can reach us here.

If any of these notes or ideas resonate with your team, we're always open to a conversation.