Test the Harness, Not the Model¶
A live demonstration can succeed twice and fail on the third run. It proves that a model answered once. It does not prove that the loop executes a tool, records the result, emits terminal events, or cleans up after failure.
Axio’s scripted transport makes those contracts deterministic.
Outcome¶
One fast agent-loop check proves the complete request sequence: model request, tool call, tool result, second model request, final text, and session end.
Fast Track¶
Replace the provider transport with
StubTransport.Script one tool-use response and one final text response.
Collect the event stream and assert behavior, not unstable model prose.
Keep live-provider and Docker checks as a smaller integration suite.
Hands-on delta¶
1. Prove one complete agent turn¶
The handler is real. Only the model boundary is scripted. This focused example does not reconstruct the final deployment harness:
async def read_document(path: str) -> str:
"""Return one document from the public repository."""
return f"{path}: Build an agent harness with Axio."
async def verify_complete_turn() -> None:
transport = StubTransport(
[
make_tool_use_response(
"read_document",
tool_input={"path": "README.md"},
),
make_text_response("README.md describes the agent harness."),
]
)
context = MemoryContextStore()
agent = Agent(
system="Use repository tools before reporting document content.",
transport=transport,
tools=[Tool(name="read_document", handler=read_document)],
)
events = []
stream = agent.run_stream("Summarize README.md.", context)
try:
async for event in stream:
events.append(event)
finally:
await stream.aclose()
results = [event for event in events if isinstance(event, ToolResult)]
iterations = [event for event in events if isinstance(event, IterationEnd)]
endings = [event for event in events if isinstance(event, SessionEndEvent)]
final_text = "".join(
event.delta for event in events if isinstance(event, TextDelta)
)
assert len(results) == 1
assert results[0].content == "README.md: Build an agent harness with Axio."
assert results[0].is_error is False
assert final_text == "README.md describes the agent harness."
assert len(iterations) == 2
assert len(endings) == 1
assert endings[0].stop_reason is StopReason.end_turn
history = await context.get_history()
assert history[0].role == "user"
assert history[-1].role == "assistant"
The check does not assert whether a provider chooses the right tool from natural language. That behavior needs a small model evaluation. This example owns the harness contract after a tool call has been emitted.
Earlier lessons verify the guard, SQLite, compaction, Docker binding, session registry, MCP lifecycle, and event adapter as separate contracts. Keep each failure close to the boundary that must handle it.
2. Build the failure matrix¶
Add one focused check for each boundary that can fail:
Failure |
Required evidence |
|---|---|
malformed tool JSON |
error |
handler exception |
|
guard denial |
handler does not run; denial is observable |
provider exception |
|
client cancellation |
stream closes; no orphaned persistent tool request |
concurrent sessions |
different contexts progress independently |
same-session overlap |
the turn lock prevents interleaved history |
sandbox startup failure |
partial resources close through |
wire encoding |
every public event produces the documented envelope |
Use Testing for the complete helper API and more failure examples.
3. Separate deterministic checks from evaluations¶
These checks answer different questions:
- Unit and harness checks
Does the system validate, dispatch, persist, stream, and clean up correctly?
- Model evaluations
Does a selected model choose the intended tool and complete realistic tasks?
- Integration checks
Do the provider, SQLite database, MCP server, Docker daemon, and delivery framework work together in the target environment?
Do not make the fast suite depend on API credentials or a Docker daemon. Run a smaller integration suite where those services are available.
Try It¶
Run uv run python examples/tutorial/test_the_harness.py from the repository
root. Then run your project’s lint and type checks.
Done when¶
[ ] The scripted turn proves one tool result and exactly two iterations.
[ ] The stream ends with one
SessionEndEvent.[ ] Each important failure boundary has one focused check.
[ ] Model evaluations do not replace deterministic harness checks.
[ ] External-service checks are isolated from the fast suite.
Capstone¶
Run the finished harness against a real project task such as:
Read the repository documentation, find one failing test, make the smallest correct change, run the relevant verification, and report the evidence.
Observe the event stream, token growth, guard decisions, session ownership, and sandbox lifetime. When behavior fails, add the smallest deterministic check that reproduces the harness failure before changing the implementation.
You now have the important boundary Axio is designed to provide: a provider- independent loop inside an application-owned, observable, testable harness.
Continue with Core Concepts, How-To Guides, or the API Reference.