Harness Engineering for Long-Running Agent Skills
Published by O'Reilly Media, Inc.
Compile task-specific instructions and maintain explicit execution state
What you’ll learn and how you can apply it
- Treat agent skills as versioned units of capability that move through discovery, selection, compilation, execution, evaluation, and repair
- Build a runtime compiler that selects the appropriate skill and assembles only the instructions required for the current task
- Design explicit execution-state schemas and merge rules for long-running skills without continuously replaying an expanding transcript
- Combine deterministic validation with model-based evaluation using relative judging, judge-verifier patterns, and judge pools
- Implement targeted repair loops that isolate a failed output section and retry it using only the relevant output, state, and skill rules
- Trace skill versions, compiled instruction surfaces, model bindings, state updates, evaluation results, and repairs so that agent behavior can be reproduced and debugged
Course description
Agent skills package reusable procedures, specialized knowledge, and supporting resources, but defining a SKILL.md file is only the beginning. The surrounding harness has to decide which skill applies, assemble the instructions required for the current task, and preserve the right execution state as work continues. Without these mechanisms, agents accumulate unnecessary context, replay increasingly long transcripts, and become more expensive and difficult to test as tasks grow, while the overall task accuracy degrades.
In this two-hour hands-on course with Nicole Koenigstein, you’ll build a runtime compiler that assembles task-specific instructions and maintains explicit execution state for long-running skills without continuously replaying an expanding transcript. You’ll combine deterministic checks with model-based evaluation, including relative judging, judge-verifier patterns, and judge pools. You’ll learn how your harness can isolate a section when an output violates a skill rule and how to perform a targeted repair using only the relevant output and related skill instructions. You’ll leave with working Python code and a reusable approach for building versioned, traceable, and reliable long-running agent capabilities.
This live event is for you because...
- You’re a senior software engineer, AI platform engineer, ML or MLOps engineer, solution architect, or AI infrastructure architect building or maintaining LLM-powered agents.
- You already understand foundational agent patterns and want to engineer the harness around the model rather than focusing on prompts or framework abstractions.
- You want to use skills as modular, versioned, and testable units of agent capability.
- You’re working with long-running agent tasks and need a more reliable approach than repeatedly passing complete execution transcripts back to the model.
- You want to improve the observability, reproducibility, cost, and maintainability of agent execution.
- You need more reliable evaluation and repair strategies for requirements that can’t be verified through deterministic checks alone.
Prerequisites
- A conceptual understanding of LLMs and foundational agent patterns, including tool calling, agent loops, structured state, and workflow orchestration
- Experience interacting with an LLM through an API and working with prompts, model parameters, and structured responses
- Intermediate Python programming skills and a basic familiarity with JSON or YAML and schema-based validation
- Familiarity with agent skills or filesystem-based instruction artifacts (helpful but not required)
- No experience with a specific agent framework required
Recommended preparation:
- Set up a Python 3.12 environment (best in Google Colab)
- Install dependencies from the provided GitHub repository (link to come)
- Obtain an API key for OpenRouter and a free account for either Langfuse or LangSmith
Recommended follow-up:
- Read Harness Engineering (book)
- Read AI Agents: The Definitive Guide (book)
Schedule
The time frames are only estimates and may vary according to how the class is progressing.
Compiling and evaluating task-specific agent skills (60 minutes)
- Presentation: How skills function as versioned units of agent capability and move through discovery, selection, compilation, execution, evaluation, and repair; how to combine deterministic checks with semantic evaluation, use separate models for generation and judging, and choose between relative judging, a judge-verifier, and a judge pool
- Demo: Building a runtime compiler that selects the appropriate skill, resolves task-specific variables, assembles only the applicable instructions, and records the skill version and compiled runtime surface used for an execution
- Hands-on exercises: Compile a skill for a predefined task, generate several candidate outputs, and compare deterministic validation with relative model-based judging; inspect how the selected skill version, compiled rules, model bindings, and evaluation results are captured in the execution trace
- Q&A
- Break
Maintaining state and repairing long-running skill execution (60 minutes)
- Presentation: Why long-running skills require explicit execution state rather than an expanding transcript; how structured state supports reliable evaluation and correction over time; how targeted retries isolate a failed output section and combine it with only the skill rules required for its repair
- Demo: Extending the skill harness with an explicit state schema and comparing full transcript replay, a fixed context window, and structured state over the same workload; how a judge detects a localized semantic failure, after which the harness extracts the affected section and performs a focused retry
- Hands-on exercise: Process a stream containing updates and delayed corrections, evaluate the resulting deliverable, and repair one failed section without regenerating the complete output; compare context growth, retained information, evaluation results, retry scope, and cost
- Q&A
Your Instructor
Nicole Koenigstein
Nicole Koenigstein is an AI researcher and practitioner in agentic systems, working across research, consulting, teaching, and direct system implementation to build reliable, production-ready AI systems. Her work focuses on multi-agent architectures, evaluation, safety, and long-term system behavior.
She served as an external evaluator for a European Commission AI Grand Challenge and has advised IOSCO on generative AI in regulated environments. She also serves on advisory boards for leading AI and quantitative finance conferences and regularly delivers invited talks and technical workshops across academia, industry, and international events.
Nicole is the author of AI Agents: The Definitive Guide, Transformers: The Definitive Guide—Applications Beyond NLP, and the forthcoming Harness Engineering, all for O’Reilly, as well as Math for Machine Learning and Transformers in Action for Manning.