---
title: Harness Engineering for Production AI Agents (October 2026)
description: OpenAI's 11 February 2026 Codex experiment merged about 1,500 pull requests. AgentCore Harness, GA 17 June 2026, still defaults to 75 iterations and a 3,600-second timeout — the harness sets those bounds, not the model.
url: https://www.factualminds.com/blog/harness-engineering-production-ai-agents-2026/
datePublished: 2026-10-11T00:00:00.000Z
dateModified: 2026-10-11T00:00:00.000Z
author: palaniappan-p
category: AI Agents
tags: bedrock-agentcore, strands, amazon-bedrock, ai-agents
---

# Harness Engineering for Production AI Agents (October 2026)

> OpenAI's 11 February 2026 Codex experiment merged about 1,500 pull requests. AgentCore Harness, GA 17 June 2026, still defaults to 75 iterations and a 3,600-second timeout — the harness sets those bounds, not the model.

**Checked 11 October 2026.** An AI agent in production is a model inside a system that decides what it can see, what it can call, how it recovers, how the work is checked, and how you score the result. A stronger model does not install that system.

[OpenAI's harness engineering note](https://openai.com/index/harness-engineering/) (Ryan Lopopolo, 11 February 2026) describes one internal product built with Codex: on the order of a million lines, about **1,500 merged pull requests**, three engineers later seven, about **3.5 pull requests per engineer per day**, and an estimate of roughly one tenth the time of writing it by hand. Humans steered. The repository, the linters, the app boot, and the review loop did the rest. Those figures are OpenAI's account of that experiment, not an industry benchmark and not a FactualMinds result.

On AWS, [AgentCore Harness reached general availability on 17 June 2026](https://aws.amazon.com/about-aws/whats-new/2026/06/amazon-bedrock-agentcore-harness-generally-available/). The [get-started guide](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/harness-get-started.html) still says that if you omit a model, the harness uses **Anthropic Claude Sonnet 4.6 on Amazon Bedrock**, and that a single invocation stops at **75 iterations** or **3,600 seconds** unless you set a tighter ceiling ([operations](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/harness-operations.html)). The [Strands harness quickstart](https://strandsagents.com/docs/user-guide/harness/quickstart/), read the same day, names a different default: **Claude Opus 5** on Bedrock, plus shell, file, and web tools. Same word. Different control plane.

> **Cited published figures (not a FactualMinds run)** — OpenAI, 11 February 2026: about 1,500 merged PRs, about 3.5 PRs per engineer per day, `AGENTS.md` kept to roughly 100 lines, single runs of up to six hours. Source: [Harness engineering](https://openai.com/index/harness-engineering/).

> **First-party signals we reuse (not a new agent engagement)** — Gateway server-side tools cut median tool round-trip **~180 ms to ~95 ms** on a B2B CRM assistant (12 tools, ~8k turns/day) — [Gateway post](/blog/amazon-bedrock-agentcore-gateway-server-side-tool-execution-2026/). Support-style AgentCore at **50K sessions/mo is ~$791/mo** platform plus model — [decision guide](/blog/aws-bedrock-agentcore-vs-amazon-q-enterprise-decision-guide-2026/). There are still **zero** published FactualMinds case studies of a production AI agent.

> **Reproduce this** — Clone the folder under [`examples/architecture-blog-2026/harness-engineering/`](https://www.factualminds.com/examples/architecture-blog-2026/harness-engineering/README.md). `python3 -m py_compile harness_policy.py agentcore_invoke_sketch.py strands_harness_sketch.py` then `python3 harness_policy.py`. Expected lines include `approval:blocked`, `retry:reconcile_before_retry`, and `verify_ok:pass`.

**Opinionated take:** write the task contract, the tool allowlist, and one deterministic check in application code before you pick a host. Use [AgentCore Harness](/blog/production-ai-agents-aws-agentcore-harness-strands-2026/) when you want AWS to run a single-agent loop and you will still author IAM, Cedar, and evals. Use [Strands harness](/blog/strands-harness-lower-token-cost-2026/) for a general-purpose coding or research agent you can run locally, and do not point its default shell at orders or payments. Trade-off: you give up the one-line local loop when you move a commerce workflow onto Gateway. You keep a boundary the model cannot talk its way through.

## Four meanings of "harness"

Teams collide these. Separate them before you compare prices.

**Harness engineering** is the discipline. You design the task contract, the context the model sees, the tools it may call, the memory it may keep, the checks that accept the work, the permissions around side effects, and the record you use later. You can build that with application code, an agent framework, a managed service, or a mix.

**Amazon Bedrock AgentCore Harness** is a managed loop on [AgentCore Runtime](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/harness-vs-runtime.html). You declare model, instructions, tools, memory, and limits. AWS runs the loop inside a Firecracker microVM per session. CloudTrail records harness operations under `AWS::BedrockAgentCore::Runtime`. There is no separate harness charge. Framework choice and graph-style workflows are not on Harness; export to code when you need them.

**Strands harness** is an open-source agent you install and run. Python: `pip install strands-harness` and `create_harness()` from `strands_harness` (Python 3.10+). TypeScript: `npm install @strands-agents/harness` and `createHarness()` from `@strands-agents/harness` (Node.js 20+). The CLI on-ramp is `npm install -g @strands-agents/cli`. It returns a plain Strands `Agent`. Providers documented on 11 October 2026: Amazon Bedrock, Anthropic, OpenAI, Google, Ollama, and LiteLLM. It is built on the Strands Harness SDK, which you can use directly when every default is wrong. It does not grant AWS microVMs, Gateway, or Cedar.

**A custom Strands agent on AgentCore Runtime** is your process, packaged for Runtime (`BedrockAgentCoreApp`, ARM64 container or CodeZip). Runtime isolates the session. Memory, Gateway, Browser, Code Interpreter, and outbound Identity are calls you make. Exporting a managed harness (`agentcore export harness`) generates Strands Python so you can take that path without rewriting the tool list by hand. Claude Agent SDK export is still listed as coming soon ([export](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/harness-export.html)).

![Side-by-side diagram: AgentCore Harness is configuration inside AWS, while a Strands agent on Runtime owns the loop and only gets Gateway if the code calls it](../../assets/images/blog/harness-engineering-production-ai-agents-2026-compare.webp)

Managed execution and a local `create_harness()` can sit in one program over time. The overlap that hurts is two memories, two tool catalogs, or two trace streams for the same task. Pick one system of record for the contract, one for tool authorization, and one for the run log.

## Where each approach fits

| | Custom application harness | Strands harness SDK and CLI | AgentCore Harness | Strands on AgentCore Runtime |
| --- | --- | --- | --- | --- |
| Who controls the loop | Your code | Strands defaults, every one overridable | AWS, from your config | Your code on a microVM |
| Effort to first useful run | Highest | Low for a research or coding agent | Low for a single-domain agent | Medium after you have a container and a role |
| Customization | Total | High, including throwing defaults away | Config, plus Lambda hooks where the grid says mixed | Total, inside the Runtime contract |
| Who deploys | You | You (laptop, Lambda, Fargate, EKS, or a container host) | AWS, from `CreateHarness` or `agentcore deploy` | You deploy the artifact; AWS runs the microVM |
| Memory | Whatever you store | Files under `./.agent` unless you set `session.dir` and `memory.dir` | AgentCore Memory: short-term, and long-term strategies (semantic, summarization, user preference, episodic) with an actor id | You call Memory, or you bring your own |
| Observability | Your traces | Strands telemetry, then wherever you export it | AgentCore traces into CloudWatch | Runtime plumbing plus what you emit |
| Security you still own | All of it | Sandbox, tool permissions, untrusted input | IAM, Cedar, prompt checks, skill sources, evals | The same, plus the framework |
| Cost meters | Model, your compute, your logs | Model, your compute, your logs | Runtime CPU and memory, model, and any Memory, Gateway, Browser, Code Interpreter, or CloudWatch you use | Same AWS meters if you call those services, plus your framework's token use |
| Best fit | A narrow workflow you already run in code | A coding or research agent with a human nearby | First production agent on AWS whose topology fits one loop | Multi-step code you will not express as harness config |

Employee-only knowledge work has a third door. Compare [Quick Suite and AgentCore](/blog/agentcore-harness-strands-vs-amazon-quick-chat-agents-flows-2026/) before you install a shell-capable agent for people who already have seats.

## Why a stronger model still fails in production

The failure is usually outside the weights.

The task was vague, so the agent edited a workflow file. The tool list included `deploy` because the prototype needed it. The context window filled with old tool output and the instruction that mattered was summarized away. A retry repeated a refund whose first response had timed out. The model graded its own patch and called it done. A silent model update changed the tool-call shape and nobody pinned the id.

OpenAI's write-up is explicit about the early weeks: progress was slow because the environment was underspecified, and the fix was almost never "try harder." They made the app, the logs, and the metrics legible, and they kept rules in linters so the agent could not forget them. They also still ask Codex to review its own change locally and then request other reviewers. Self-review is one step in that loop. It is not the acceptance test.

> **What broke** — OpenAI, 11 February 2026, on their own repository: one large `AGENTS.md` crowded out the task, went stale, and could not be checked mechanically. They cut it to a map of about 100 lines and moved the system of record into `docs/`, with linters for structure and freshness. Source: [Harness engineering](https://openai.com/index/harness-engineering/).

I looked for a study that says self-evaluating agents approve their own output 94% of the time and a separate verifier cuts that to 61%. I could not find a primary source that measures those two rates on agent acceptance. Nearby papers use 94 and 61 for other quantities (math accuracy, iterative self-correction, or "everything looks good" feedback). Those are not this claim. Leave the pair out. Use a test.

The same standard applies to "40% of the budget went to replaying old context" and to "optimize when the accepted-change rate falls below 50%." I could not trace either as a measured result. Acceptance rate is a team choice. It depends on the value of a good change, the cost of a bad one, and the cost of review. A 50% line is not a standard.

## Ten responsibilities

![Reference architecture: application entry, task contract, context builder, model loop, tool boundary, memory, policy, independent verification, telemetry, and recovery](../../assets/images/blog/harness-engineering-production-ai-agents-2026-architecture.webp)

Each box in that figure is a responsibility. Prompt text can remind the model. The stop has to live in code, IAM, a sandbox, a policy engine, or a runtime limit.

### 1. Task contract

Write the objective, the success checks, the allowed surface, and the prohibited operations before the model plans. The plan is a proposal. The contract is the constraint. Record preconditions (clean git status, order id that exists, customer matched to the tenant) and the evidence required before anyone marks the task complete. If the contract cannot name a file, a tool, or a business rule, the task is not ready for an agent.

### 2. Tool orchestration

Give each agent a small set of tools with typed arguments. Keep planning and execution apart when a wrong call is expensive. Deterministic code should own predictable steps: tax, eligibility, a status lookup with one id. Timeouts, retry caps, and idempotency keys belong on the tool runner. Delegation to a subagent pays off when the extra coordination is cheaper than stuffing the whole job into one context. The model must not be able to emit a tool call that skips validation. On AgentCore Harness, AWS rejects a caller-supplied `toolUse` block in the final `InvokeHarness` message, so a client cannot name `shell` and have it dispatched directly ([security](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/harness-security.html)). Your own Runtime entrypoint does not get that rejection unless you add it.

### 3. Context

Build the window from the instructions, files, and tool results this task needs. Do not paste the repository or the full history into every call. When you summarize, keep the contract, the safety rules, and the open decisions verbatim. Mark facts as authoritative, uncertain, stale, or untrusted. Issue bodies, web pages, and tool output are untrusted even when the user is trusted.

### 4. State and memory

Keep these apart:

| Kind | Lives for | Correction |
| --- | --- | --- |
| Task state | This attempt | Discard on cancel |
| Session history | This conversation | Compact, but keep the contract |
| Durable memory | Across sessions | Provenance, expiry, delete |
| Project instructions | Until a human changes them | Review, do not let the agent edit its own guard file unnoticed |
| Workflow state | Until the business process ends | Required to resume after a crash |
| Business data | In the system of record | Read it again. Do not treat memory as the order |

Tenant and user isolation, retention, and deletion are properties of the store. On AgentCore, per-user memory uses an actor id, and per-user credential brokering through Identity's token vault requires inbound OAuth. SigV4 does not propagate per-user identity into downstream tools today; AWS says that support is planned ([security](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/harness-security.html)). Storing a transcript is not the same as extracting a preference. See [agent memory for commerce](/blog/ai-agent-memory-ecommerce-2026/).

### 5. Verification

Run checks the model does not grade: unit and integration tests, typecheck, lint, policy, schema, migration dry-runs, commerce rules (order state, amount, owner), and a person for high-impact changes. Agent eval suites catch regressions in behavior; they do not replace those checks. A second model is another opinion. It shares failure modes with the first when the rubric is vague. Evaluations belong next to the deterministic suite, which is the subject of the [ecommerce regression post](/blog/ai-agent-evaluation-regression-testing-ecommerce-2026/).

### 6. Guardrails

Four layers, from the practical split we use in reviews:

| Layer | Examples | Enforced by |
| --- | --- | --- |
| Behavioral | Output shape, refused topics, task limits | Prompt plus a schema check on the final message. [Bedrock Guardrails](/blog/how-to-set-up-amazon-bedrock-guardrails-production/) if you want a managed filter |
| Data | PII redaction, retention, tenant isolation, what may be shown | Application code and the trace sink. A prompt does not redact logs |
| Tool and action | Allowlist, argument constraints, draft-then-commit, approval | Tool runner, IAM, Gateway Cedar |
| Operational | Iteration cap, deadline, token budget, concurrency, cancellation | Runtime settings or your loop counter |

### 7. Observability

Record run id, task id, model id and configuration, tool name, redacted arguments, result or error, latency, retries, context truncation, memory hits, verifier results, human decisions, token counts, and the final outcome. Skip raw prompts, full customer payloads, and secrets. The observability post is the longer version of what to dashboard. Traces that contain everything are a second data store you now have to protect.

### 8. Routing

Send cheap, fast models at classification and extraction when you have eval evidence they hold. Send a stronger model at work where a wrong answer is expensive. Also account for latency, price, and whether that model id is enabled in the account. Do not freeze a nickname ranking from last quarter. A routing model is the wrong tool for "if the intent is WISMO, call `getOrder`." That is a table.

AgentCore documents mid-session provider switching that rebuilds history into the next provider's format. Use it for a measured shift, then re-run evals. The Strands `effort` argument accepts `auto` (default), `low`, `medium`, `high`, and `off`, mapped per provider ([model configuration](https://strandsagents.com/docs/user-guide/harness/configure/model/)).

### 9. Reviewed feedback

When the verifier fails:

1. Store the failure and the evidence.
2. Name the cause. A bad test is a different fix from a bad tool schema.
3. Propose one narrow change to instructions, tools, permissions, tests, context selection, routing, or recovery.
4. A person reviews that change.
5. Add or update a regression test.
6. Run the harness on a representative set, not only the one case.
7. Release in a controlled way and watch both misses and false rejects.

Do not append every rejection to `CONSTRAINTS.md`. A wrong verifier, or a rule that blocks the common case, makes the system less useful. This is harness maintenance. It is not fine-tuning unless you actually train weights.

### 10. Recovery

Handle a half-finished task, a process restart, a tool error, a contradictory check, stale memory, a broken task record, and an upstream outage. Cap retries. Open a circuit when the dependency is down. Escalate to a person with the evidence attached.

Before a retry, answer one question: did the first call change remote state? If you do not know, read the system of record. If it changed, do not send the write again. If it did not, retry with the same idempotency key.

## Context rot, and what to persist

Context rot is the practical failure of a long window: the instruction you needed is still "in" the transcript and the model no longer behaves as if it were. OpenAI's failed encyclopedia file is the clean public example. Compaction and summarization, which both AgentCore (`sliding_window` or `summarization`) and Strands (offload bulky tool results, compact as the window fills) will do for you, can drop the same line unless you pin it outside the summary.

Keep a reproducible record of the context and the configuration used for a decision you may have to defend: contract version, model id, tool catalog version, and which documents were retrieved. You do not need the full prompt body to do that.

Strands sessions default to `./.agent/sessions` with a generated id. Pass `session={"id": "..."}` to resume, and on anything but a laptop set `session={"dir": ...}` and `memory={"dir": ...}` to durable storage ([production](https://strandsagents.com/docs/user-guide/harness/production/)). AgentCore Memory is a different store: short-term events and long-term records with the strategies listed in the [harness versus Runtime grid](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/harness-vs-runtime.html). Runtime's own microVM filesystem dies with the session. Memory is how you keep something after that.

## Tools, permissions, and prompt injection

![Sequence from a model tool proposal through schema validation, authorization, optional human approval, execution, and an independent check, with a deny path](../../assets/images/blog/harness-engineering-production-ai-agents-2026-sequence.webp)

Treat retrieved content as untrusted. [OpenAI's prompt-injection note](https://openai.com/safety/prompt-injections/) describes the pattern: a third party hides instructions in content the agent was asked to read. I am not repeating an unverified story about a crafted GitHub issue executing a script. The control is the same either way. The reader of untrusted text should not hold the credentials that perform the side effect. AWS is explicit that Harness does not inspect prompt meaning, and that skills fetched from S3 or Git are treated as trusted input. If callers can set `skills` or `additionalParams` on `InvokeHarness`, they can point the agent at another repository or pass provider fields through, including endpoint overrides on some providers. Strip those fields at your application edge ([security](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/harness-security.html)).

Strands ships shell, file (`read`, `write`, `edit`), and web tools, and a `programmatic_tool_caller` that runs model-authored code in Monty. Monty isolates that code. It does not isolate the tools the code calls. The production guide says to sandbox and to use interventions before untrusted input, or to drop the tool. The default system prompt says to confirm before irreversible actions. That sentence is not an authorization boundary.

On AWS, the Gateway path is the one we use for anything that writes. The [Gateway post](/blog/amazon-bedrock-agentcore-gateway-server-side-tool-execution-2026/) measured tool round-trip, not safety. Safety is Cedar plus the absence of the dangerous tool. `InvokeHarness` also requires `bedrock-agentcore:InvokeAgentRuntime` on the harness ARN, because the harness is a Runtime resource. The sample execution role is yours to narrow.

## Verification and the feedback loop

![Feedback loop from an independent check through a named cause, a reviewed harness change, a regression test, and a controlled release, separate from model training](../../assets/images/blog/harness-engineering-production-ai-agents-2026-feedback.webp)

A coding agent that says "tests passed" has not passed tests until your runner's exit code says so. A support agent that says "refund issued" has not issued a refund until the order system says so on a fresh read.

Wire the seven steps above into the place you already review changes. The artifact's `verify_coding_task()` returns failure codes. An empty list is the only pass. The model is not in that function.

## What you pay for

Token price and run price are different bills.

There is **no separate AgentCore Harness charge**. The [pricing page](https://aws.amazon.com/bedrock/agentcore/pricing/), checked 11 October 2026, bills the capabilities you use. For a harness session that is at least Runtime, plus the model, plus CloudWatch if you keep the traces. Add Memory, Gateway, Policy, Browser, Code Interpreter, or Web Search only when that architecture calls them. A managed harness is not automatically cheaper than a loop you run yourself. You trade engineering time and a platform meter against tokens and a container you operate. Model it on the [AgentCore pricing calculator](/tools/amazon-bedrock-agentcore-pricing-calculator/) and read the [twelve-component pricing note](/blog/amazon-bedrock-agentcore-pricing-12-components/).

Runtime microVMs, consumption rates on that page:

| Meter | Rate |
| --- | --- |
| v1 CPU | $0.0895 per vCPU-hour |
| v1 memory | $0.00945 per GB-hour |
| v2 CPU, consumption | $0.1276 per vCPU-hour |
| v2 memory, consumption | $0.0169 per GB-hour |

CPU can drop during model and tool wait when nothing else is using it. Memory stays billable while the session is up. Runtime v2 reclaims idle memory after 120 seconds. There is a 128 MB minimum and a 1-second minimum. The worked example on the same page, a support-shaped session with 60 seconds of 1 vCPU active time and their stated memory profile, comes to about **$0.006703 of Runtime compute per session** before model tokens, Gateway, Memory, and logs. That is AWS's example, not a measurement from this site. A harness session also counts against Runtime quotas, including **2 vCPU and 8 GB** per session ([limits](https://docs.aws.amazon.com/bedrock-agentcore/latest/devguide/bedrock-agentcore-limits.html)).

Other meters, same page, include only what you enable:

- Browser and Code Interpreter: $0.0895 per vCPU-hour and $0.00945 per GB-hour, 128 MB minimum. Leaving them on for ordinary chat is a common way to inflate the platform line. The [August harness post](/blog/production-ai-agents-aws-agentcore-harness-strands-2026/) already flagged that pattern.
- Gateway: $0.005 per 1,000 API invocations (list, invoke, ping). Search is $0.025 per 1,000. Tool indexing is $0.02 per 100 tools per month.
- Policy: $0.000025 per authorization request. Natural-language authoring is $0.13 per 1,000 input tokens, separate from the per-request charge.
- Web Search: $7 per 1,000 queries.
- Memory, as of 6 October 2026: short-term ingestion $1.00 per GB (each event billed at least 12 KB and at most 64 KB), retrieval $0.20 per GB, storage $0.10 per GB-month on actual size. Long-term built-in storage $0.75 per 1,000 records per month. Long-term retrieval $0.50 per 1,000 retrievals.
- Identity: no extra charge when used through Runtime or Gateway.
- Observability: CloudWatch ingestion, storage, and query. Do not copy a CloudWatch unit price out of an AWS example without checking the CloudWatch page.

What inflates tokens, on any host: a large system prompt, unused tool definitions, retrieved memory, full history, repeated calls, and retries. What inflates Runtime: a long `maxLifetime` (default 8 hours) and a long idle timeout (default 15 minutes) on sessions that sit warm.

Strands token comparisons, including the 21 September 2026 Harbor numbers, stay in the [Strands post](/blog/strands-harness-lower-token-cost-2026/). They are not an AgentCore invoice.

**Illustrative accounting only.** Suppose 10 attempts cost C in total and 4 are accepted by the independent check. Cost per attempt is C/10. Cost per accepted change is C/4. Those denominators answer different questions. Nothing here sets C, and a 4-of-10 outcome is not a target.

## AgentCore Harness, as documented

CLI (Node.js 20+), from the current get-started page. The older `@aws/agentcore@preview` line in the launch blog is not the command that page shows now.

```bash
npm install -g @aws/agentcore
agentcore create --name coding-task-demo --model-provider bedrock
agentcore deploy
```

Python sketch, boto3, `DRY_RUN=1` by default. Poll `get_harness` until `READY` in real code. Full file: [`agentcore_invoke_sketch.py`](https://www.factualminds.com/examples/architecture-blog-2026/harness-engineering/agentcore_invoke_sketch.py).

```python
# Python 3.10+, boto3, region where Harness is GA. DRY_RUN prints; it does not call AWS.
control = boto3.client("bedrock-agentcore-control", region_name="us-west-2")
created = control.create_harness(
    harnessName="coding-task-demo",
    executionRoleArn="arn:aws:iam::123456789012:role/HarnessExecutionRole",
)
# Poll get_harness until status is READY. The get-started page calls this field arn.
harness_arn = created.get("arn") or created.get("harnessArn")
client = boto3.client("bedrock-agentcore", region_name="us-west-2")
response = client.invoke_harness(
    harnessArn=harness_arn,
    runtimeSessionId=str(uuid.uuid4()),  # 36 chars; minimum is 33
    messages=[{"role": "user", "content": [{"text": "Summarize git status. Do not edit files."}]}],
    maxIterations=8,
    timeoutSeconds=900,
)
```

Stream events follow `messageStart`, `contentBlock*`, `messageStop`. `stopReason` can be `end_turn`, `tool_use`, `max_iterations_exceeded`, `timeout_exceeded`, `max_output_tokens_exceeded`, or `hook_stopped` when a `before_invocation` or `after_tool_call` Lambda hook returns deny. Hooks are mixed: the configuration is yours, the Lambda is code.

Lifecycle defaults if you omit them: `maxIterations` 75, `timeoutSeconds` 3600, `maxTokens` unset, `idleRuntimeSessionTimeout` 900, `maxLifetime` 28800. Inline client-side tools need your code. Built-in shell and `file_operations`, skills, Gateway, Browser, and Code Interpreter are configuration. Skills are trusted content. Review them before you attach a Git URL.

Export when config is not enough:

```bash
agentcore export harness --name coding-task-demo --build CodeZip
```

Then read `EXPORT_NOTES.md`. The generated agent is a Runtime agent. Deploying it does not delete the harness you exported from. Turn one of them off for that task so you do not operate two loops.

## Strands harness, as documented

This API is not the block above.

```python
# Python 3.10+, pip install strands-harness. Confirm the model id in your account.
from strands_harness import create_harness

agent = create_harness(
    model="global.anthropic.claude-opus-5",
    session={"dir": "/var/lib/agent/sessions", "id": "api-design"},
    memory={"dir": "/var/lib/agent/memory"},
)
```

The model page shows `global.anthropic.claude-opus-5` as a Bedrock id and `anthropic/claude-sonnet-5` as a provider string. The quickstart's unnamed default is Claude Opus 5 on Bedrock. Pin one id. `effort="high"` is optional. The sketch that matches this call is [`strands_harness_sketch.py`](https://www.factualminds.com/examples/architecture-blog-2026/harness-engineering/strands_harness_sketch.py). It does not import the AgentCore clients.

Out of the box, Strands also gives you a tuned system prompt (explore, then act, confirm before irreversible steps, verify before finishing), context offload, prompt caching on Bedrock and Anthropic, long-term memory, a `generalist` subagent, a `todos` tracker, and Agent Skills when they are present ([harness overview](https://strandsagents.com/docs/user-guide/harness/)). Production hosting targets in their deploy guide include Lambda, Fargate, EKS, and AgentCore. "Listed as a host" means the process can run there. It does not mean CreateHarness.

An earlier post on this site recorded Claude Opus 4.8 as the quickstart default in September 2026. The quickstart checked on 11 October 2026 names Claude Opus 5. Re-read that page when you pin.

## Pattern A: a coding agent

The contract is a data type, not a prompt.

```python
# Python 3.10+, stdlib. From harness_policy.py in the artifact folder.
contract = TaskContract(
    task_id="codemod-retry-copy",
    objective="Update the retry message in src/utils/retry.ts and keep tests green.",
    allowed_paths=("src/utils/retry.ts", "src/utils/retry.test.ts"),
    prohibited_operations=("git_push", "deploy", "add_dependency"),
    required_checks=("unit", "typecheck"),
    max_iterations=8,
    deadline_seconds=900,
)
```

Walk the run in this order.

1. Read the project instructions and this contract. The instructions are a map. The contract is the scope.
2. Record `git status` before any edit. A dirty tree is a precondition failure, not something the agent should tidy by force.
3. Ask for a plan that names the files it expects to touch. Reject the plan when a path sits outside `allowed_paths`.
4. Execute with a command allowlist: `git status`, `git diff`, the unit test, the typecheck. `git push`, dependency installs, and workflow edits are absent from the runner, not merely discouraged.
5. Run the checks yourself. `verify_coding_task()` fails the run on an empty diff, a missing check, or a path outside scope. In the sample, a `package.json` edit returns `path_outside_scope:package.json`.
6. Compare the diff to the objective. A green test on the wrong file is a failure.
7. Write a `RunRecord`: run id, task id, model id, iterations, token counts, tool failures, outcome. Leave the prompt body out.
8. Stop when iterations, the deadline, or an unknown remote side effect says stop. `require_approval("git_push", None)` returns `blocked`.

`retry_decision(mutated_remote_state=None, ...)` returns `reconcile_before_retry`. That is the rule for a `git` or API call whose HTTP status you never saw.

## Pattern B: a commerce support agent

The customer asks for a refund. The agent may draft the reply. It may not be the thing that authorizes the money.

Identity first. The session carries a tenant id and a customer id from your authenticator, not from the message text. `getOrder` accepts those three ids and returns the minimum fields the reply needs: status, total, and a shipment state. Street address and full payment numbers stay in the order system. This is the same boundary as the [support agent](/blog/ai-customer-support-agent-ecommerce-2026/) and the [store security](/blog/secure-ai-agents-ecommerce-store-2026/) posts.

The refund tool is not on the tool list for that turn. The agent can submit a proposal: order id, amount, reason. `require_approval("refund", approval_id)` returns `blocked` until a person in the [human-review path](/blog/human-in-the-loop-ai-agents-ecommerce-2026/) writes `approval_id`. After that, Gateway Cedar (or your API authorizer) checks the amount and the order state again. The model does not mint the approval id. When the write returns, read the order once more and attach that status to the trace. A timeout on the write is `mutated_remote_state is None`: reconcile, do not send a second refund.

Cancel, inventory adjustment, and any reply that would reveal another customer's data use the same gate. A prompt that says the agent is a careful associate is the behavioral layer. It is not this gate.

## A custom policy, and the line that enforces it

[`application-policy.yaml`](https://www.factualminds.com/examples/architecture-blog-2026/harness-engineering/application-policy.yaml) is a review document. The header says it is not an AgentCore schema and not a Strands config file. `harness_policy.py` does not load it. The [enforcement map](https://www.factualminds.com/examples/architecture-blog-2026/harness-engineering/enforcement-map.md) is the part to keep open while you implement.

Abbreviated:

```yaml
# Custom application policy. Not applied by CreateHarness or create_harness().
schema: factualminds.application-harness-policy/v1
scope:
  allowed_paths:
    - src/utils/retry.ts
    - src/utils/retry.test.ts
iterations:
  max_iterations: 8
  deadline_seconds: 900
cost:
  stop_after_model_calls: 12
approvals:
  required_operations: [git_push, deploy, refund, cancel_order, disclose_pii]
verification:
  required_checks: [unit, typecheck]
  agent_self_approval_counts: false
```

| YAML field | What stops the action |
| --- | --- |
| `allowed_paths` | `verify_coding_task()`, or a sandbox that cannot read the rest of the disk |
| `max_iterations` / `deadline_seconds` | Your loop, or AgentCore `maxIterations` and `timeoutSeconds` |
| `stop_after_model_calls` | Code that refuses the next model call. Not AWS Budgets |
| `required_operations` | `require_approval()`, and IAM or Cedar on the tools you exposed |
| `required_checks` | The test runner's exit code |
| `agent_self_approval_counts: false` | You never add a "mark complete" tool the model can call alone |

A `CONSTRAINTS.md` file is a reasonable place to keep rules the reviewer accepted. It becomes real when a linter or the tool runner fails the build for a violation. Until then it is a document the model can ignore, the same way OpenAI's oversized `AGENTS.md` was a document the agent could not reliably follow.

## How to tell whether it is working

Define the denominator before you celebrate a rate.

| Metric | Denominator |
| --- | --- |
| Task success | Tasks whose independent checks passed, over tasks you intended the agent to attempt |
| First-pass verification | Attempts that passed with no harness retry and no human edit |
| Regression rate | Previously passing cases that fail after a harness change |
| Acceptance rate | Changes a reviewer or a check accepts, over changes proposed |
| Human correction time | Minutes a person spends per accepted task |
| Tool failure and retry rate | Tool calls that error or repeat, over tool calls |
| Escalation rate | Tasks handed to a person, over tasks started |
| Latency | p50, p95, p99 of the whole task, not of one model call |
| Cost per successful task | Model plus platform plus logs, over tasks that passed |
| Cost per accepted change | That same total, over changes that were accepted |
| Policy violations caught | Denies that matched a real forbidden action |
| False-positive rejects | Denies or check failures on work that should have shipped |
| Cost of failed work | Spend on attempts that were abandoned or repeated |

A high first-pass rate on an easy set is a weak signal. Include cases the agent should refuse, cases with missing data, and cases where the tool fails halfway. Inject those on a schedule. Roll out a harness change to a slice of tasks and watch false rejects, not only the happy path. The evals post covers how a green dashboard hides that.

## Mistakes that show up in the first month

- The tool catalog from the prototype, including shell and a write API, is still attached.
- Browser or Code Interpreter stays on for turns that only needed a Gateway read.
- Session files sit on ephemeral disk, so a new task "forgets" and the model re-reads everything.
- Compaction drops the contract. The agent expands scope.
- Retries replay a payment or a refund.
- The agent can edit the file that contains its own allowlist.
- Two harnesses run the same task after an export, and the traces disagree.
- Every verifier complaint becomes a new bullet in the prompt, and the prompt gets worse.
- Model id is "latest", so a provider-side change looks like a random quality drop.
- AWS Budgets is the only spend control, and the bill arrives anyway.

The [Monday checklist](https://www.factualminds.com/examples/architecture-blog-2026/harness-engineering/monday-checklist.md) is the short version of the fixes.

## What to Do This Week

1. Pick one workflow. Write the contract: objective, allowed tools or paths, prohibited operations, evidence of done.
2. Remove every tool the workflow does not need. Refunds, deploys, and pushes require an approval id your code checks.
3. Add one deterministic check and fail the task when it fails. Do not ask the model whether it passed.
4. Log model id, tool name, redacted arguments, tokens, and outcome. Set an iteration cap and a deadline in the runner or on `InvokeHarness`.
5. If you are on AgentCore, confirm the execution role, turn Browser off, and put writes behind Gateway. If you are on Strands, point session and memory at durable storage and sandbox the shell before any untrusted text.
6. Price the mix on the [calculator](/tools/amazon-bedrock-agentcore-pricing-calculator/). Keep the OpenAI and Strands benchmark numbers in their own columns.

If you want a second pair of eyes on the contract and the tool boundary, start with the [AI agents hub](/ai-agents/) or [contact the team](/contact-us/). We will not pretend a harness diagram is a finished commerce agent. Gateway policy, identity, and the eval file are still yours. The [production AgentCore guide](/blog/amazon-bedrock-agentcore-production/) is the AWS map. This post is the discipline that map sits inside.

## What This Post Doesn't Cover

- A worked fine-tune or a change to model weights. The feedback loop here edits the harness.
- Current Bedrock token prices. They move. Read the model page for the id you pin.
- A full Cedar syntax tutorial. The Gateway and [store security](/blog/secure-ai-agents-ecommerce-store-2026/) posts carry that.
- Multi-agent topologies (Graph, Swarm, Workflow). Use the [August Strands 1.0 post](/blog/production-ai-agents-aws-agentcore-harness-strands-2026/) when you outgrow one loop.
- Any claim that FactualMinds has shipped a production agent for a named client. We have not published one.

## How the Harness engineering stays bounded

Put the model inside controls for observation, action, recovery, verification, and evaluation.

- **Maturity:** Level 2 — reference architecture
- **Risk:** high
- **Oversight:** approval
- **Starts when:** A team is deciding how a production agent is allowed to act.
- **Tools:** A write goes through a router, a permission check, a policy check, and a budget check. The model does not commit it. An approval token lives in tool context, not in the user message.
- **Stops for a person:** Prompt text is not authorization. Refunds, pushes, deploys, and other writes stay behind a named gate. The model does not approve its own output.
- **Checked by:** An independent check runs before the task is accepted. A second model is not the first check.
- **Untrusted data:** Tool results, tickets, issues, and retrieved memory that try to change the task or the policy.
- **If it fails:** If the Harness engineering stops, name the reason: completed, budget_exceeded, timed_out, cancelled, guardrail_blocked, approval_required, tool_failure, verification_failed, or partial_completion. Retry a timeout at most twice. Back off on rate limits. After repeated verification failure, escalate. Stop when the budget is exhausted or a permission is denied. No unbounded loop. Prompt text is not authorization. Refunds, pushes, deploys, and other writes stay behind a named gate. The model does not approve its own output.
- **Lifecycle:** Explore, then act, then confirm, then verify. Explore gathers the context for the Harness engineering. Act stays reversible. Confirm stops before an irreversible step. Verify is a separate check of the outcome.

Reference architecture, not a published production deployment. The shared contract is the [AWS store-agent architecture](/blog/aws-ai-agents-for-ecommerce-factualminds-2026/#operating-contract).

## FAQ

### When should I not start on Amazon Bedrock AgentCore Harness?
Skip the managed harness when you already need a graph, a workflow with hard step order, bidirectional streaming, or a framework other than the one AgentCore exports. The harness-versus-Runtime page lists framework choice and non-agent-loop patterns as unsupported on Harness. Use Runtime and your own code, or keep the agent off AgentCore for a single Converse call.

### What could go wrong if the system prompt is the only guardrail?
The model can still call any tool you attached. AWS documents that AgentCore Harness validates the shape of InvokeHarness, not the meaning of the prompt. A line that says do not refund does not remove the refund tool. Put refunds, cancels, inventory writes, and private-data reads behind IAM, Gateway Cedar, or application code, and require a human approval id before those tools exist for the call.

### Is Strands harness the same product as AgentCore Harness?
No. AgentCore Harness is the managed AWS loop (CreateHarness and InvokeHarness, GA 17 June 2026). Strands harness is the open-source package strands-harness and @strands-agents/harness. create_harness() returns a Strands Agent you run yourself. Hosting that process on AgentCore Runtime does not attach Gateway, Identity, or Policy unless your code calls them.

### What could go wrong if an agent approves its own output?
The same model that wrote the change is a poor judge of it. OpenAI's harness write-up still has Codex review its own diff, then request separate reviewers, and it relies on linters and CI for rules that must hold. Prefer a test, a schema check, or a second read of the system of record. A second model can help and can still be wrong.

### Does AWS Budgets stop a runaway agent immediately?
No. Budgets and CloudWatch alarms notify you. They are not an instantaneous cap on the next model call. AgentCore documents per-invoke limits you can set: maxIterations (default 75), timeoutSeconds (default 3600), and maxTokens (unset unless you set it). Idle session timeout defaults to 900 seconds and max lifetime to 28800 seconds. An application counter that refuses another call is the other hard stop.

### Should every verifier failure become a new permanent rule?
No. Record the failure, name the cause, review a narrow harness change, add a regression test, and watch false rejects after release. A wrong verifier or an over-tight rule reduces the tasks the agent can finish. This loop changes instructions, tools, tests, or permissions. It does not train model weights.

---

*Source: https://www.factualminds.com/blog/harness-engineering-production-ai-agents-2026/*
