Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162660 stories
·
33 followers

Cloud Performance Testing with Postman: From Laptop to CI

1 Share

Cloud performance testing gives you what a laptop can’t: load spread across many machines, a clean baseline, and results your whole team can see. You set the virtual user (VU) count. Postman handles the fleet, the tuning, and the merging, and you get one set of numbers you can trust.

Most teams still start performance testing on a laptop, and that’s fine. This post covers what the cloud adds, why that matters more as AI agents become some of your heaviest API consumers, and how to put it in your pipeline.

Why cloud performance testing beats your laptop

On a local run, the machine running the performance test becomes part of the test. Your CPU, Wi-Fi, VPN, and uplink all sit between the load and the API. When numbers look bad, you can’t tell if the API is slow or your laptop is. A laptop also runs out of room: pushing more load means raising OS limits and tuning TCP, which can have side effects and may need admin rights you don’t have. And two runs on two days, on two networks, aren’t a fair comparison.

A cloud run sends traffic from Postman’s infrastructure instead. You get:

  • A clean baseline. The load generator isn’t competing with your email client.
  • A real network path. Requests arrive over the public internet, the way remote users and agents reach you.
  • Repeatability. The same runner setup every time, so a change in results means a change in your API.
  • Scale. Postman’s cloud ramps to millions of virtual users. A one-hour soak at 2,000 concurrent users isn’t a stretch.
  • Shared results. Every run, including runs started from CI, lands in Postman with its history and trend. Your team looks at the same numbers.

Load comes from the region of your Postman account: the US by default, or the EU on EU Data Residency plans.

What Postman’s cloud takes care of

Cloud runs use the same collection you run locally. The difference is the work you no longer do. (For how it works under the hood, see Performance Testing in the Cloud with Postman: What We Built.) Here’s what you own if you build the setup yourself, and what you don’t with Postman.

Scaling out On your own: you size the machines, set up a fleet, split the virtual users across them, and get every machine to start at the same time. With Postman: you set --vu-count. Postman picks the machines and spreads the load across them.

Machine tuning On your own: every load generator needs higher file-descriptor and port limits and tuned TCP settings. Those go into the image and have to stay the same across the fleet. On a laptop or a locked-down machine, you may not be allowed to change them at all. With Postman: the load generators come ready.

Merging results On your own: you collect the output from every machine and combine it. Correct percentiles take extra work, and you need a dashboard to show them. With Postman: metrics from every machine come back as one run, with live and final latency percentiles, throughput, and error rate.

One test, one tool On your own: load scripts live in a separate tool, apart from the API collections your team already maintains. With Postman: the collection, auth flow, dataset, and pass condition you use for local performance testing run in the cloud and from CI. No rewrite. A few things stay local-only: data files, mTLS client certificates, and Local Vault secrets. For cloud runs, use a dataset and Shared Vault secrets instead.

Setup and teardown On your own: you write separate scripts to seed and reset test data around the load test. With Postman: --setup-collection runs once before the load starts and --teardown-collection runs once after it ends, whatever the outcome. They run as their own stages, outside the load phase.

Fixed egress IPs On your own: you run NAT gateways or reserved IPs, and you update allowlists whenever the fleet changes. With Postman: on Enterprise, --runner postman-cloud-static-ip sends load from a dedicated cluster with a fixed egress IP that you allowlist once.

Nothing to keep running On your own: you patch images, upgrade the test runner, shut machines down after each test, and pay for capacity that sits idle between tests. With Postman: you pay only for the VU-hours you use.

Why AI agents make performance testing matter more

Human traffic has a shape. People click, read, and click again. Agent traffic doesn’t.

Agents are fast becoming a main audience for APIs, and they don’t behave like people. An agent working on one task can call your API many times in a row. It can call several endpoints at once. It doesn’t take breaks, and it doesn’t keep office hours. One user request can turn into dozens of API calls, and many users doing this at once can produce load your capacity planning never saw.

That makes the shape of the load as important as its size. Postman performance testing has four load profiles (–load-profile), and each one maps to a question about agent traffic:

  • fixed: Can the API hold steady concurrency over time? The VU count stays constant. This is your baseline for agents that run all day.
  • ramp-up: Where does it start to degrade as more agents come online? VUs climb from 25% to 100%, then hold. Use it to find the knee in the curve before adoption finds it for you.
  • spike: What happens when a lot of agents start at once, such as a scheduled job firing or a launch going out? VUs start at 10%, jump to 100%, then drop back to 10%.
  • peak: Can it survive staying near maximum load? VUs climb from 20% to 100%, hold, then come back down to 20%. Agents don’t go home, so peak load can last for hours.

Each profile is a fixed pattern across the run’s duration, which you set in minutes. The profiles are VU-based: rps is something you can gate on, not a rate you drive load at. If your team talks about load tests and stress tests, Performance testing vs. load testing vs. stress testing explains how the terms relate.

Agents also don’t forgive slow or flaky responses. A person waits. An agent times out, retries, or picks another path. So your p95 and your error rate become product behavior.

Why does the cloud matter here? Agent load can be large, steady, and long. A laptop hits its own limits first. A cloud run spreads the load across machines and keeps going for as long as the test needs.

One thing to plan for. Cloud runs come from Postman’s IP ranges. If your API allowlists by IP, or rate limits per IP, use the static-IP runner (Enterprise) and allowlist it. Otherwise your performance test measures your firewall instead of your API.

Make performance tests realistic with datasets

Sending the same request a thousand times is misleading: caches answer fast and hot rows stay in memory.

Attach a dataset so each virtual user sends different real inputs. Datasets work the same way wherever the run happens, which makes them the right choice for cloud and CI runs. You can connect a dataset to a live MySQL, PostgreSQL, or SQL Server database on Team and Enterprise plans, so you’re not exporting stale CSVs. Custom JDBC sources need Enterprise. Use a test or staging copy, or a read-only account with anonymized data.

For agent traffic, include what agents actually send: long prompts, large payloads, unusual parameter combinations, and a few malformed requests.

How do I add performance testing to CI with a p95 gate?

A performance test you run by hand gets run before big launches and not much else. A test in your pipeline runs every time.

Cloud runs fit CI well. A CI runner is small and shared, so it can’t generate much load, and its numbers move with whatever else is running on it. A cloud run takes the load off the runner. The build only waits for the result.

If you haven’t set up the CLI yet, start with Working with the Postman CLI. Then sign in with an API key, add a pass condition, and the test becomes a gate:

postman login --with-api-key "$POSTMAN_API_KEY"

postman performance run <collectionId> \ --runner postman-cloud \ --load-profile ramp-up \ --vu-count 50 \ --duration 10 \ --pass-if "less_than(p95, 500)" \ --output ndjson

This runs the collection from Postman’s cloud with 50 virtual users for 10 minutes. It exits with code 1, and fails the job, if p95 latency goes over 500 ms. The check happens after the run, so a bad build still sends its full load before the gate fails. Start with a low VU count and a short duration on anything live.

You can gate on avg, p90, p95, p99, error_rate, or rps, with less_than, less_than_eq, greater_than, or greater_than_eq. –output ndjson streams results as newline-delimited JSON, which suits CI logs and coding agents where the terminal dashboard can’t render.

When local performance testing is still the right call

Local performance testing is fast and free on every plan, and it’s the better choice when you’re building the test itself, when the API is on localhost or behind a VPN or firewall, when it needs mTLS, or when you’re checking a dev build for obvious problems like a slow query or a memory leak. Don’t pay for more than you need.

Start local. Move to the cloud when the answer needs to be one you’d bet a launch on.

When is a script-first tool a better fit?

Postman is the shortest path when your API tests already live in Postman collections and you want performance testing without rewriting them. A code-first load tool may fit better if you need:

  • Second-by-second ramp shaping, such as 50 to 3,000 clients in exactly 30 seconds.
  • An arrival-rate model that drives a target requests-per-second instead of a VU count.
  • The load test itself, not just the pipeline step, to be a plain script file in your repo.

Try it

Take a collection you already run locally. Attach a dataset. Run it once locally and once in the cloud, then compare the two. The gap is what your laptop was hiding. Once you trust the numbers, add the --pass-if line to your pipeline.

FAQ

Can Postman run load tests from the cloud, not just my machine? Yes. postman performance run <collectionId> --runner postman-cloud sends load from Postman’s managed cloud infrastructure. Without --runner, load comes from the machine that runs the command.

Does Postman performance testing need the desktop app? No. The Postman CLI runs performance tests headlessly in any CI system. The desktop app is one way to configure a test, not a requirement for running it.

How many virtual users can Postman generate? Postman’s cloud ramps to millions of virtual users. Cloud runs need at least 10. A local run is limited by the machine it runs on.

Can I fail a CI build when p95 latency is too high? Yes. Add --pass-if "less_than(p95, 500)". The command exits with code 1 when the condition isn’t met, which fails the CI job. This works for local and cloud runs.

Is Postman performance testing free? Local performance runs are unlimited on every plan, including Free. Cloud runs cost $0.04 per VU-hour on Solo, Team, and Enterprise, with pay-as-you-go turned on.

Can I allowlist Postman’s load generators in my firewall? Yes, on Enterprise. –runner postman-cloud-static-ip sends load from a static-IP cluster in your account’s region (US, or EU on EU Data Residency plans).

Where do cloud runs send load from? From your Postman account’s region: the US by default, or the EU on EU Data Residency plans. Each run uses one runner, so you can’t combine local and cloud load in the same run.

The post Cloud Performance Testing with Postman: From Laptop to CI appeared first on Postman Blog.

Read the whole story
alvinashcraft
5 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Steer Your Agents Until They Steer Themselves

1 Share

Over the last year and a half I have been working side by side with Postman’s CTO, Ankit Sobti, on the agent platform we use to deploy all of our internal agents as we transform the company to be AI-native. This is a follow-up to his post on the Agentic OS we built, which covers the business need that led us here, an architectural overview, and the technical precedent that shaped our thinking. In this article, I will discuss the problem space surrounding building agents for enterprise, in particular for non-verifiable qualitative domains (not coding/math), our approach to solving them, and the open problems that remain. I hope this serves as a useful guide for high-agency engineers or engineering leaders hoping to steer such an effort.

Context

I came to Postman after working on a startup. There is a stereotype that technical founders would rather code their way through problems than talk to customers. Unfortunately, these stereotypes can occasionally be true. With no salespeople or budget, I too chased the dream of a product that grew itself through scrappy LLM-augmented outreach tools, but learned the hard way that no amount of automated go-to-market (GTM) can grow something that doesn’t have a market. In coming to Postman, I actively sought the opposite problem. Postman had already reached 40 million users through product-led growth. So, a primary objective was to just convert an increasing subset of those existing users into paying customers by delivering real enterprise value, and to keep shaping the product around what they told us about their new AI needs.

I helped found our forward deployed engineering (FDE) team, wrote its first playbook and shipped our first major deployments, which led to significant million-dollar expansions in annual recurring revenue. Although we were able to get to a great outcome with these individual customers, the process behind it could not scale. We relied on our executives’ intuition to pick the right accounts. We could only read so much of the calls, emails, product telemetry and support tickets on each one, and even what we did read missed context that lived only in the account team’s heads, outside our systems of record. Filling the gaps meant more discovery on the customer’s schedule, and nothing we learned reached the sellers, so even a problem we had already solved came back to us.

What we needed was a GTM machine that relentlessly compounds. Everything we know about a customer would live in our systems of record, and the machine would read all of it and tell our people which accounts to approach, with what, and what to ask the core product team or FDE to change in our product. Anything our team had built once would become something sellers could sell on their own, which would leave FDE free for the last mile of novel engineering. Every time our people acted on what the machine told them, the result would become new data, so its next recommendation would be better informed. It also had to be cheap enough to run on every account every week, and keep working as the models and the data underneath it changed.

So I tried automated go-to-market again. This time I was armed with the luxury of a product that had already found its market. My startup-background-induced naiveté about accelerating a field that has always run on intuition and relationships at enterprise scale was shielded by our developer tool founders, who shared my rose-tinted view that we should be AI-pilled and apply agentic thinking outside software. Having learned my lesson, I wanted the best of both people and AI, and that came down to a maxim.

Abstract

Human at every surface, machine on the analysis behind it

Fig. 1: A ring of people forms the only boundary a customer touches, and inside it a dense green network of analysis.

Maximally human on external communication. Maximally AI on internal analysis. That should be the self-improvement loop of any enterprise.

Let’s start with the external half. As much as possible, every customer should feel that their problem was heard by a person who cares, and that the solution was built for their use case. Our largest accounts have a median of 45 stakeholders named on their deals, and the median relationship has run about three years, so keeping those relationships strong is one of the best uses of our people’s time. AI can handle the quick things, like a support answer at midnight or an automated onboarding flow. But an AI avatar on a sales call, or a bulk-generated cold email that blindly guesses at a customer’s problems, cheapens the brand, and that is not what being AI-pilled means.

The internal half is a great place to be AI-pilled. You want decisions made at a velocity, throughput, and accuracy no individual can mentally accomplish, and the reason is that it is biologically impossible. Research on working memory finds that we can actively juggle only about four chunks (meaningfully grouped pieces of information) at once (Cowan, 2001), where in our case 1 call would roughly be 1 chunk. Even at one large account for one year, there can be hundreds of calls and thousands of emails. To make matters worse, the highest leverage insights are often thematic across accounts and it can get up to {calls, emails, telemetry, support, docs, etc.} × years × accounts. This is terabytes of data that no person can hold while trying to draw the highest leverage line of reasoning across all of these data points.

Accounts across, years down, one route through

Fig. 2: Each cell is one account’s year of records. The green path links the data points the agent reasons across on its way to an answer, and the small box shows how little of it one person can hold at once.

Take a seller who wants to reach out to each of the customers in their book of accounts that would benefit from something we launched recently. Doing it well means reading up on each account, and with a couple hundred accounts each there is not enough time in the week. Timing matters here, because the frontier moves quickly and features go stale. An agent can go through every account for them, pull out what it finds that matters for each one, and bring the strategies our best sellers have already succeeded with to all of them. The seller can then write to every customer and tailor the message to the right stakeholder with exactly how it would benefit them.

As we grow more confident in the strategies our best sellers use, we can hand more of that work to the agent to run autonomously. The questions sellers ask over and over become skills, the skills they run every week become scheduled reports, and eventually the agent keeps its own memory of each account, asking the questions it deems important to meeting our KPIs.

To climb toward an agent that does more on its own, we can take a page from how software handled layers of abstraction. Engineers once wrote machine instructions by hand, until an assembler took over that translation. Compilers then let them write in languages people could read, like C, and object-oriented languages like Java let them bundle data with the code that works on it, so they could build large systems out of reusable parts. This ladder worked because every rung had both generation and verification. Generation meant a tool could produce the work at the new level. Verification meant engineers could check that it was right. A compiler gives the same output every time it gets the same code and refuses to build code with whole classes of mistakes, so engineers could trust each rung and build the next one on top of it.

Qualitative domains had neither automated generation nor verification before AI. Enterprise GTM is one such qualitative domain. The answers to questions like “What is the optimal strategy to sell customer X on product feature Y?” are open-ended and have no true right answer. Previously, you could not automatically generate an answer for the next move on an account without a human Account Executive in the loop. Now, an agent with tools generates an answer based on the subset of data it finds during its execution time. However, it is still hard to know whether the agent reached the optimal subset given the global corpus. Did it find that one call where they discussed their new AI strategy? Did it look at the one data point in the product telemetry where it shows they actually were leveraging your new feature but in an unconventional way? Verification remains an open problem. We have not solved it, but every approach that has worked for us starts the same way, by breaking it into smaller checkable pieces.

Generation lengthens the ladder, verification makes a rung hold

Fig. 3: Each braced software rung pairs with a check, from the hardware up to tests, while the rungs above the line, like the next move on an account, stop short unchecked.

I decided the generation half was promising enough to build on, and started with a GTM agent that answers questions about our customers from the thousands of calls we have recorded with them. Like any conventional agent, it began as a model, a system prompt defining its persona, and tools that return data, all running in a harness that accepts a question and loops until it has an answer. From there it grew into an accelerated agent deployment platform that now enables agents to be quickly spun up for any end user across the company. Since late May, nearly 400 employees have used them, from sellers and executives to product managers, marketers and engineers, through Slack, the web, MCP and the command line, both for questions they ask and for proactive reports and alerts that come to them. The coaching agents alone reach nearly all of our sellers, about 220 people a month, and have been asked almost 10,000 questions. Monthly questions grew 76% from July to September, with 80% retention, and a strong power law in usage with 1 in 10 having asked more than 100.

Spinning agents up quickly and serving answers turned out to be the easy part of that platform. Aside from verification to ensure we were improving the agent over time, a significant hurdle was getting the right people working on the parts of the agent they were best suited to. Every team has its own goals, incentives, definition of a good answer and view on who should see which data, and each is right about its own, so we were never going to write their agents for them from the outside. What makes the platform work is that each team writes only the part it understands, and everything else lives in a shared substrate we maintain for them. The agent becomes a configuration of a shared thing instead of a new thing, so every improvement to the substrate lands on every agent standing on it. That is what lets the GTM machine compound.

One substrate, molded per team

Fig. 4: Four team agents differ in shape and in their prompt, tools, data, guardrails and access settings, yet all stand on one shared substrate that improves beneath them.

1. Everybody Builds Their Own Agent

Agent building inside enterprises has reached a fervor we have not seen before, because anyone who can write instructions and connect a data source can now build an agent. Even within a single team, several people build agents that share a function, end users, and data sources. Nobody wants their agent to just sit on their laptop, so they rush to serve it to the others. Unfortunately, everyone else has the same idea and serves theirs back. Every one of these efforts stands up its own pipeline over the same data, so each new agent leaves another copy of your customer records somewhere with less oversight than the last, and those pipelines go ungoverned the moment access expires or a dependency moves, which is when the agent starts answering confidently out of data that stopped updating weeks ago. There are two common attempts at solving this, and both appear to work at first.

The first way is that the domain team or individual builds it themselves. Our sellers did, through Claude Desktop, wiring up connectors by hand. That gets you running in an afternoon, and the person doing it understands their own workflow better than any platform team will. What you cannot do is scale it to an entire team and globally steer it. Outcomes varied wildly from one seller to the next, because each seller wrote different instructions and connected a different set of sources, so two sellers asking about the same account could get answers drawn from dramatically different data. When an executive wanted every agent working the new sales play, there was nowhere to put the play. It lived in a deck and the agents lived in individual sessions, so none of them ever saw it, and the machine we set out to build depends on exactly that kind of shared context. The same gap shows up in how they read data at face value. If the customer said the product was too expensive, the agent told the seller to offer a discount. That is rarely the right answer, since any customer of any product in any industry could ask for a discount, whether or not the product serves them. What the seller actually needs might be a nudge to re-qualify the champion, or to re-anchor on the outcome the customer is buying. Those are strategies our senior sellers and executives have worked out, and a personal agent has no way to receive them.

Every seller's agent, wired by hand

Fig. 5: Six sellers have each wired a different mix of sources into their agent, and the new play sits in a deck none of the agents can reach.

The second way is to hand the work to a forward deployed or applied AI team. These are skilled software and systems engineers who have learned to work with AI. Put them together as a small team and they can handcraft one agent per internal use case. That solves steering, because those engineers own the single copy of each agent deployed to a team, so one change reaches everyone using it. What it does not solve is anything after the first version. Each agent is bespoke, so there is nobody to hand it back to, and it stays with the people who built it. Every one shipped becomes something somebody keeps alive, and making the same patch to each agent they shipped is not the highest leverage use of a capable engineer’s time.

One play, added to every agent by hand

Fig. 6: Engineers add the same new play to four bespoke agents one at a time, and the fourth, negotiation prep, is still waiting for its copy.

What became obvious once a few of these sat side by side is that they were variations of the same bounded list. An agent is a system prompt, a subset of the tools and data sources you have already provisioned, a model and a budget for how long it is allowed to think, an execution horizon that says whether it answers when asked or runs on a schedule, a set of guardrails, and a rule about who may invoke it and which fields it is allowed to read. And this list can always be expanded as you uncover new primitives you want to bring to all of your agents.

What we need is for the platform to be configurable so that each of the relevant parties in the company can write the component of the agent that pertains to them and leave the rest to the others. Access control, the hard deterministic rules governing who can reach which agent and which data fields, is best controlled by IT. The system prompt is best written by the domain team, since they know their own workflow and what they want the agent to do on their behalf. The guardrails, the soft steering markdowns that keep the agent away from the poor behavior you only learn about by deploying it, are best written by the manager. Instead of one of these teams writing every component, or an outside engineering team guessing at all of them, each component gets decided by the people closest to it. The agent platform team can then just work on improving the overall substrate.

The ideal construction

Fig. 7: Each bear built by a single team fails in its own way, and the combined bear takes access control, prompt, runtime and guardrails from the team closest to each.

2. Deploying Every Agent from a Monorepo

In order to have catch-alls for functionality the platform did not support out of the box, we let agent developers write custom code on top of the agent they were serving. That code turned out to be the signal we needed. When our engineers found themselves writing patches with the same shape for the third or fourth time, the shape got promoted: it became a field any agent could declare in its manifest, with no code at all.

Every agent is written to a single monorepo that deploys when a PR merges. An agent is a directory. Everything specific to that agent lives inside it. Two folders sit outside that everything shares: the globally provisioned tool set, which each manifest names a subset of, and the guardrails, which name the agents they apply to. The appendix shows an example of each of these pieces.

What got promoted

Fig. 8: Apart from its SKILL.md, every setting on deal-coach-pro is a choice from a fixed set, including 52 of the 54 tools and 32 of the 45 guardrails.

When it comes time for deployment, we wanted a mechanism that enabled agent developers to quickly deploy their agent without taking down a shared substrate. Astro, our agent deployment and governance platform, made this very convenient. The shape of it is ordinary Docker. Every agent builds from the same Dockerfile with the same build context, the whole backend, so COPY . . bakes every skill into the image. Each agent then carries a manifest, a short YAML file that names that Dockerfile and declares the environment variables its container needs. One of those variables is the name of the skill directory to serve. The adapter reads it at boot and serves only that one, which means the codebase is the artifact and the variable is the pointer: every pod runs the same tree with a different path lit up inside it.

One agent's path through everything available

Fig. 9: SKILL_NAME lights up the deal-coach/ directory and the shared files, tools and guardrails it names, and every gray entry ships in the same image without being loaded at runtime.

3. Getting Long-Horizon Results at Runtime Speed

One thing held across all of our agents: the longer we let a run go, the more tokens we let it burn, and the more data we gave it access to, the better it performed. This is perhaps a microcosm of Richard Sutton’s Bitter Lesson, which holds that massive computation dwarfs last-mile algorithmic iteration. The constraint with building runtime agents like the above is that users expect an answer inside a reasonable window, about two minutes.

What we tuned against what we spent

Fig. 10: Inside the two-minute window hand tuning gives the better answer, but past the crossover more tokens and longer runs keep improving it long after tuning has flattened.

You can give them the option to configure a job that runs much longer, ten minutes or more, and we did, but that carries its own set of problems. One, if everyone starts launching computationally intensive jobs at will, token spend goes out of control. Two, you stop your users from getting deeply nested in a workflow and reaching the higher order insight. When responses come back quickly they keep going down the rabbit hole, and your overall usage of the tool grows exponentially. This can also be read through Jevons paradox, which says that the cheaper you make something the more of it people use, except that inside an enterprise the thing your employees are trying to conserve is often not cost but their own time. So if they believe it is saving them real time, and more of it is at their fingertips, they will use it more, which raises their overall AI consumption and drives your transformation.

We needed to find a way of providing the benefits of longer running jobs while making the runtime inference fast and cheap. We need to find a way for an AI agent to figure out which pieces of information are the highest leverage for an account and refresh those as new information comes in every week. “Highest leverage” is a vague term that is specific to your company’s needs so to filter for this we would need another probabilistic agent steered by prompts to filter the high volume of incoming information. I will discuss two approaches I took here: a harness that allows an agent to recursively in parallel spawn sub agents so that it can better research an account, and the second being Anthropic’s first party dreaming primitive. The former was hand-rolled before the release of the latter but both have since proven to have their benefits.

The sources are given, the filter is the product

Fig. 11: Streams of data arrive all week, and one agent, set by its model, prompt, tools and token budget, keeps the few items that fill the account document.

The first in-house method is what we call the research pipeline and it runs once a week for 8 hours. Each account research process begins with a deterministic set of API calls that pulls context from each of our data APIs which took place after last week’s run into a shared store. This includes the Salesforce record, its opportunities, its account plan and its contact roles, along with transcripts from every Zoom call, product telemetry and support tickets. From here, in batches of 4, 24 sub agents are spawned that use this store as a seed to decide what threads they would like to further investigate. What they are looking for is fixed by the sub agent in advance and it is informed by what they see in the seed, what agents from previous batches already returned, and core pillars the account team cares about (renewal risk, expansion path, stakeholder map, implementation blockers, recent customer sentiment). Each of these sub agents reasons over the tool calls it is making until they find what they were looking for or the token budget runs out.

The path each one takes through the data closely resembles a graph traversal, because at every turn the agent can move vertically, deeper into the same source, or horizontally, across to a different one. For example, it might be reading a chunk of a Zoom call where the customer says a feature no longer works for them, and go vertical, back through earlier calls to see whether they turned it off, or horizontal, out to the product telemetry to see whether they are using it at all. It ends up making a combination of horizontal and vertical turns, which resembles an A* search, except that the heuristic is the model’s probabilistic evaluation of whether a tool call will yield alpha, whereas A* would use a deterministic function. At the end, every agent returns its findings to a more intelligent, large context model that synthesizes the data points and writes a detailed account dossier. This dossier is stored in a relational database and also embedded into a vector database. It is available as a tool at runtime and the agent is told to prefer it in the tool definition and system prompt since if it contains the answer to the question being asked then further tokens don’t need to be burned querying the raw data source.

Twenty-four threads, and two directions each

Fig. 12: Researching sub agents walk the seeded sources, stepping deeper or across at each turn, and their findings feed one synthesis that writes the dossier.

While the research pipeline is our best attempt at giving agents the highest leverage insights at runtime, Anthropic released Dreaming as a primitive for curating an agent’s memory directly. Rolling your own keeps the cultivated memory model-agnostic. Using a black-box dreaming API makes sense for the opposite reason. The model provider is the one most likely to know what harness curates information so that their own model recalls it the way you want at runtime.

Given that, I set up a second pipeline in parallel, on Dreams. To understand Anthropic’s Dreaming you have to start with their managed Memory Stores.

Memory stores exist to solve the ephemeral nature of agent sessions. An agent makes its tool calls, reasons over them, the container goes away, and everything it worked out is lost. To avoid that you need a persistence layer, and the obvious ones are the databases you already run. That is what we did for the research pipeline: the dossiers go into SQL and the transcripts behind them into a vector index. The problem with relational and embedding-based retrieval is that the agent is only ever shown the matching subset. If you want Acme you can run SELECT * FROM dossiers WHERE account = ‘acme’. If you want to know where onboarding is going badly across the book you can embed “onboarding issues” and take the nearest neighbors. Both work, and both are narrow by design: the agent reasons over the matches and never learns what else was there.

What a query reaches

Fig. 13: The SQL query returns one Acme row and the embedding search four onboarding chunks, but the security review chunk holding the answer falls outside both.

Now ask why Acme’s rollout has stalled. Both queries fire and both come back full. The dossier has the onboarding tickets, the training sessions, the adoption curve flattening in March. The semantic search has a dozen accounts with the same complaints and the themes they share. There is plenty to reason over, and the agent writes a confident answer about enablement. The actual reason might be a sentence their security lead said on a call fourteen months ago: they cannot grant the directory access the rollout depends on until a review that never got scheduled. That sentence is not semantically close to “onboarding issues”, it sits outside any sensible recency window, and it was filed against a stakeholder, nowhere near the rollout. Nothing in either result set tells the agent it is missing. An agent that can see the whole corpus and navigate through it knows what it is choosing to leave out.

A memory store enables the agent to see the entire corpus. It is a collection of small text files that lives above any session, with an id of its own, and every file has a path, so a store has the shape of a directory. You attach one when you create a session, read-only or read-write, and the platform mounts it inside the container. Writes the agent makes there persist back when the session ends, and each one leaves an immutable version behind recording what changed and who changed it.

There are no memory tools. No search endpoint, no embeddings, no index. The agent runs ls and gets back every path in the store, not the subset that matched something, and it sees those paths whether or not it would have thought to ask for them. From there it navigates: grep for a term, cat a file, and decide from what it just read what to open next, using the same tools it is already well trained on from navigating code. In addition, since we don’t need to respond in a highly constrained time horizon like the runtime agent, we can afford to let it explore this data to its satisfaction.

What an agent reaches

Fig. 14: The agent lists all files, greps for rollout down to three, then reads timeline.md, which points it to security-lead.md, a file the grep never matched.

That brings us to how we used these primitives to build our account memory. The first decision is what an account has in it. We settled on one store per account: overview.md at the root, a file per person under stakeholders/, a file per opportunity under deals/, what customers actually said under quotes/, and a timeline.md of recent activity. Every file carries a provenance header naming where it came from, who asserted it and when. We seed a store from the root data sources and the output of the research pipeline, which makes every file in it a synthetic artifact we need a steered Claude to keep current. New material arrives through ingestion jobs that listen to those same sources and write it into the filesystem as a record of recent updates. This is fine for new data from the week, but if we keep piling these in, then our overall system degrades to just a store of the raw data sources. We need a weekly process to compact these updates into our high-level synthetic artifacts.

This is where dreaming comes in. A dream takes the account’s store, a model, a steering prompt called instructions, transcripts of the sessions that read it, and writes a new store. In our case those sessions are the runtime agent’s own inference, each one a different traversal of the same data. Those traversals are what let the pass decide how the data should be structured and what is worth keeping, based on the shape of the questions actually being asked. It makes the memory layer self-healing: gaps get identified, what nobody reads gets discarded, and the answers people keep walking to end up where they will be found first. A SQL or vector store will serve any model that queries it, but its schema is fixed the day you design it, and no amount of being used will change its shape.

The hard part of a memory layer is that it has to stay durable without inventing, because it is no longer surfacing raw data but syntheses of it, massaged over time. Two protections do that work. The first is a rule: the pass may not add anything that was not in its input. The second is structural: nothing it writes counts as memory until a pointer says so. A store’s id changes every pass, so the runtime never holds one. It looks up a slug and gets back whichever store is current, and that lookup is the only thing that makes a store’s data appear in the output. So a new store sits there, referenced by nothing, while it is checked: file count within forty percent of the input, nothing oversized, every view present, provenance valid throughout. If it passes, the lookup is repointed and the old id goes into history. If it fails, the lookup does not move, and last week’s store keeps answering.

One week of an account's memory

Fig. 15: Raw updates pile up all week until the weekly pass, steered by session transcripts, writes a candidate store that the slug points to only after it clears all checks.

4. The Agentic Mountain Range: Measured Agent Improvement

In order to best describe what agent building feels like, I would like you to picture yourself as a determined traveler in the foothills of a majestic but intimidating mountain range. Ranges like the Himalayas, Andes, or Zagros were so intimidating that they quite literally divided civilizations. You, on the other hand, prepare to chart your path across. As you look up at the steep slopes, you ask yourself what the best strategy is. The reality is that you are not trying to summit anything. You just want to get to the other side.

Which makes every foot of elevation a cost you would avoid if you could, and shifts a lot of the burden onto choosing: which mountains stand between you and the other side, which of them you need to cross, and in what order you take them. A party that takes the tallest peak first and dies on it has crossed nothing. You save the taller ones for later, when one of them is the only path left.

The mountain range in front of an agent team has several mountains, and those are only the ones you can see. More appear once you have covered enough ground in your own domain. A few are universal. Delivery: getting a good answer out of the system at all. Evaluation: knowing whether the answer was any good when it is hard to verify. Self-improvement: a system that suggests changes and improves itself. Coordination: an agent that knows your organization’s roles and can message the right person to do the right thing at the right time. Action: an agent that knows how to do the thing responsibly instead of only recommending it.

The range

Fig. 16: A map charting out challenging problems in agent building as mountains and a logical path of increasing resistance only when necessary.

In deterministic software, choosing the right problem carried less weight, because a bad choice announced itself. You attempt the hard problem, you fail, and the failure is checkable, so you find out early and correct course. In the probabilistic world a team can spend a year on a hard problem and whether they got anywhere at all is just their word. Did you really build a recursively self-improving agent? I guess these responses look better than before. Good job, keep going. Or maybe you did build it, and the problem runs the other way: the outputs all look alike, text and quotations, and nobody can say whether they are better than last month or worth more. So how do you make your progress clear?

In the deterministic era you knew concretely how much a change you were about to push would improve the overall system. Now that LLMs have made it probabilistic and the outputs qualitative, it is very difficult to know the impact of any given change. For this reason, software engineering in the probabilistic era has moved toward outcome engineering. This is the process of defining what you would like the outputs to look like for your coding harness, letting it make the technical decisions, scanning what comes back against your intuition for how it should respond, deploying, and then refining the system over time where it does not behave as you expect. As a result, the incentive to review every incremental change with the paranoia you once would have had is gone.

The prose diff

Fig. 17: Rewriting one instruction turns a discount plan into advice to re-anchor on the integration timeline, and nothing in either answer shows which is better.

Improving the system breaks down into climbs of different sorts. The two classic ones in computing are hill climbing and climbing the layers of abstraction. The first concentrates on one output metric, makes a change, and keeps the change if the metric moved. The second starts a system at its simplest understood form, watches it run in production, works out what you are willing to accept as fact, and then builds on top of that so the next decision gets made at a higher level. Coming back to the climb, a mountaineer drives a bolt into the face, clips the rope to it, and runs the next stretch from there. The incremental changes are the holds (or hills) you climb between one bolt and the next, and each bolt is a check you can take as fact before you build the next layer on it. What the bolt buys is that a mistake above it costs you a rope length instead of death.

Two kinds of climb

Fig. 18: Hill climbing’s scored gains shrink from +6 to +1 before the metric stalls at a peak, and the abstraction climb stacks from chat upward with a bolt at each boundary.

Hill Climbing for Agents

Hill climbing is a classic optimization technique where you look at the states adjacent to the one you are in, step to a better one, and repeat until nothing next to you is better. This means you stop on the first peak you reach, a local maximum. There are well-understood escapes from this, like sometimes accepting a worse step on purpose (see simulated annealing for arranging cells on a chip). Regardless, observe what is fundamental to the climb and the escape: a score, computed by a machine, for any given state. Coding and math are domains where this holds, and it is why the frontier labs have taken their models so far on both. Roughly, they collect a large bank of problems with strict test cases (like SWE-bench, which is real GitHub issues and the pull requests that closed them), change the model, and see whether it passes more of them. Nobody has to read the output to know, so the loop runs over terabytes, at a scale no amount of human review could keep up with.

Hill climbing for coding

Fig. 19: Test files score every answer, so each kept model passes more of the 7,200 clones, from 771 at v1 to 5,663 at v7, and the neighbors v7a and v7b score lower.

None of that holds for an agent whose output is prose about an account. Even if we keep the construction of the agent, the harness, identical, there are still infinitely many adjacent states because many of the surfaces you steer with are text. A system prompt, a markdown guardrail, a tool definition: each can be changed to any other text, so there is no bounded list of neighbors to walk. The tool definitions alone carry more than eight thousand words of English, and a good part of that is instruction written to correct something a model kept doing. The surfaces that are bounded, turn ceilings and token ceilings, multiply what is left by another order of magnitude. And then the knobs everyone reaches for first: which model you run, and which subset of the tools you hand it. Once you have landed on a configuration, what you need is a way to say whether the answer it produced was any good.

Before you can get a machine to say it, you try to codify the judgment with a person in the loop. What the work actually looks like is this. A question comes back answered badly. You change one of the surfaces above and run it again. Sometimes it is clearly better. More often it is just different, and you decide. What is the true right answer to “What is the optimal API governance strategy at Acme?” And how would you know whether the agent picked the right set of data points out of everything it could have read? So you might keep the change, you might not, and you go to the next thing. Whether it was a good change could depend on who reviewed it, how many times they ran the case to see whether it was flaky, and whether what they wrote was a band-aid or a fix. Do that for a year and it builds a sediment, a whack-a-mole game of changes made to correct behavior nobody could attribute.

So you do the obvious thing and try to do it systematically, at higher volume, by handing the scoring to a model. You take two versions of the agent with one concrete difference between them, a model change or a code change, and you pass both answers to the same question to a judge. The problem is that a judge with no tools of its own has no more idea which answer is better than an uninformed third party would. It never saw the corpus, so it cannot tell whether either answer reached the part that mattered. What it can see is the prose, so the prose is what it grades: which one reads better, which one moves between paragraphs more smoothly, which one cites more. Give the judge tools and you land in a catch-22: a judge that can go and find the best subset of the data is the agent you were trying to build, and if you had it there would be nothing left to compare.

We can use a historical analogy to understand what to do here. In the early 1700s, Antonio Stradivari, an Italian violin craftsman in Cremona, made what are considered to be the finest violins ever built, and in the three hundred years since nobody has come close. Later in life, it became apparent to Stradivari that he needed to distill his craftsmanship to others or the craft would die with him. He started with just his sons Francesco and Omobono but expanded this to other apprentices in the town to create a workshop. Stradivari was obsessive over his work and did not want to tarnish his reputation, so he needed a process for strict quality control. Like in our agent building case, the output metric of violin production is qualitative and subjective: its sound. He could not sit and listen to every single violin out the door, and the sound of the violin was only possible far too late in the process after it was carved, closed, varnished, and strung.

Nor could he hand the listening to his apprentices. Asking an apprentice whether their violin sounded as good as one of his was asking them to judge work finer than their own, with the same judgment that had fallen short in making it. That is the position of a model or LLM-as-judge today. You could have a stronger model grade two configurations of the same agent running on a weaker one. But how do you climb to the best possible agent when it already runs on the best model you have?

So there is no score, and the work goes on without one. What accumulates is a growing pile of patches, each written against a symptom, and each a bet on how one model behaves. The model updates or you change providers and every bet is open again, because the new one fails differently and the corrections you wrote now push against nothing. Worse, a patch made for a reason that has since gone away does not announce itself. Something gets fixed upstream, the symptom disappears, and the paragraph you wrote to lean against it stays in the prompt, leaning. You end up steering with controls that may already be pointing the wrong way, and nothing that would tell you.

Lucky for us, we can take inspiration from Stradivari. He needed other output metrics he could measure the quality of a craftsman by. He developed several: the body shape needed to match his hand-crafted walnut mold, the eyes of the f-holes needed to sit an exact angle apart and were measured by compass, the outline had to match his paper patterns, and the purfling, the inlaid border, had to run a constant distance from the edge. These are all hard checks that anyone in the workshop could conduct with a tool instead of his ear. None of these tools told him whether a violin would sing, but each told him something he believed needed to be true before it could. For an agent, these are verifiable sub-problems: pieces of an answer that can be deterministically checked without anyone judging the whole.

Checking a violin without an ear

Fig. 20: Body shape, outline, f-hole placement and purfling each pass a check against a workshop tool, but the sound has no tool and is left to Stradivari’s ear.

The gut reflex in agent building is one more patch, another small change to the prose, and it leaves you no higher than the last one did. Checking has no ceiling, so every time you break the work into verifiable sub-problems, you earn a way up to a harder problem. Once the agent passes them reliably, you can stand on that layer and climb to the next one. The prose changes in between are just hill climbing to a local maximum.

Hill climbing is the wrong hill to die on, and I will die on that hill.

Climbing the layers of abstraction for agents

In the case of agents it is hard to know in what order to add complexity. There are so many moving parts available to an agent developer: system prompts, context, memory, caching, persistence. Someone building from scratch is tempted to throw all of it at the wall and see what sticks, and what they end up with is an agent nobody can steer, because nobody can tell which part is responsible for anything. The order matters more than the parts.

The order has to be deliberate, and it comes from watching the people who use your agent and deciding where you are willing to give up steering by hand in exchange for leverage. For our agent the path went like this. We started with a chat agent that reached our data sources through tools, and watched which prompts people ran over and over. Those became skills, markdown files passed in at runtime, so people could skip straight to the uses we wanted to promote. Then we watched which skills got run most often, because that told us what workflows people actually had, so we let users add a cron schedule to a skill. We called the result a report, delivered to Slack or an inbox. That made visible the shape of things people want to hear about, how often, and whether they act on it when it lands. It was enough to work out what else would be worth telling them and at what cadence, which is where the proactive agent came from: one with a memory of its own, tracking what matters about each account and what something has to look like before it is worth a DM. The research pipeline and dreaming are what made that memory possible.

Each of those steps needed a bolt before we could trust the one above it. The bolts go in at the transitions between layers: a check you can clip into that tells you the layer below works as you expect, so you can keep climbing. We drove the first one in as we moved past the chat layer. Models at the time still made up numbers often enough to worry us (ARR, renewal dates, license counts), and that gave us a verifiable sub-problem: of the figures in the answer, how many never appeared in anything the tools returned while the agent explored? A high share meant the agent was likely inventing data. The same move carried up the climb. Skills exist because people ask the agent the same question over and over, which makes them a good place to measure flakiness with one number: how many tool calls a run makes. The same question should take about the same amount of digging each time. A skill that makes four calls on one run and forty on the next is not one you want on a schedule, answering for itself. For dreaming, the bolt is the pointer from section three. A new store answers nothing until it passes its checks, and until then last week’s store keeps answering.

Driving the next bolt

Fig. 21: The cliff rises in layers from chat through skills and reports to memory, with the rope clipped to a bolt at each transition and the climber drilling the next one.

Between two bolts the softer components are free to move, the prose in your prompts, skills and tool definitions, and that is where the hill climbing described above belongs. Add rigid layers without knowing why, fail to bolt them down with a check, or keep piling on prose instead, and you inch toward a slop cannon: an agent whose answers and actions nobody can explain. Research makes this easy to see: you cannot let an agent research on its own until you know it is finding the right data as it traverses, weighing relevance against recency the way you would, and that its tools are not failing quietly or returning something unexpected. You only learn that through rigorous synchronous use. Without that understanding, every call the agent makes and every answer it writes leans toward slop, and letting it run on its own compounds it. The effectiveness of these verifiable checks led us to dedicate the next mountain to increasingly comprehensive evals of non-verifiable domains.

5. Evaluating Non-Verifiable Domains

Evals are a fundamental part of agent development because they let you change the agent and be confident about what the change will do. That confidence matters more as the agent’s use grows across the enterprise, because a degraded agent does not look degraded to the people using it. It still answers with plausible facts in fluent prose, but the facts may be lower leverage than they were, and the enterprise can spend months making weaker decisions before anyone notices. Changes to an agent fall roughly into two groups: its model and its harness. A model change usually means upgrading, switching providers or moving to a cheaper model, and in every case the question is the same: what share of the performance do you gain or give up, for what share of the cost? A harness change is what your developers make as they watch the agent perform, usually adding a new layer of abstraction or hill climbing the prose on the existing ones, and the question is whether it made the agent better or worse.

The last section showed that the usual ways of answering either question, mostly borrowed from verifiable domains, fall short here: a model as judge never sees the corpus, an automated benchmark needs a right answer, and a human in the loop grades against a bar the agent is meant to far surpass. That last one is the problem the agent exists to solve. No human can hold the data behind a single account, let alone the trends across many, so tuning the agent to an expert’s answers tunes it toward answers worse than even a simple wiring of every data source could give. Verifiable domains are different. In mathematics, even the hardest problem at the International Mathematical Olympiad has a proof a mathematician can write and check, so a model’s answer can be graded.

With no answer key and no perfect answer to approach, what we have left is direction: the agent gets better as the subset of the data it surfaces out of the global superset it could reach gets better, until the returns diminish. What I propose is the breaking down of qualitative domains into verifiable sub-problems. They are how we climb, taken in concentric rings, each checking a larger part of the answer against fixed artifacts that existed before the answer did: what the tools returned, what a rule says, what a test case declares. Each gives a number you can compare before and after a change. Since nobody has to read an answer to get it, you can run as many cases as you need, each several times because probabilistic surfaces make any single run flaky, and see how often each check passes. That gives a better answer to the first question, what share of the performance a model change gains or gives up for its cost. It helps with the second too, because every harness change becomes a hypothesis you can test: try as many as you have, keep the ones that raise the pass rates, and drop or modify the rest. Our list of verifiable sub-problems keeps growing as we find more whose pass rate rises consistently when the agent gets better, so I will highlight four, in increasing size of the check, that we use on our agent and should generalize to agents in most qualitative domains.

The smallest check asks whether the evidence arrived. An agent writes its answer from whatever its tool calls brought back, and when a call fails, the turn goes on: the harness hands the model the error as the call’s result, and the answer can read just as confidently without the data as with it. “The renewal looks on track. No recent activity stands out.” could have come from a quiet quarter with no email at all, or from an email tool that failed, and nothing indicates which. How often calls fail depends on more than the state of the tools. It moves with the model, which may leave out an argument it needs, invent one that looks valid, or call the wrong tool altogether. It moves with the harness too: a reworded tool description, skill or prompt can lead the agent to call a tool the wrong way.

There are ways to make every call well formed, and none is free. A strict schema makes the model send every required argument, but a model made to fill a field it does not know can invent the value, turning a failure you can see into a wrong answer you cannot. Letting the call fail keeps it visible and leaves the model to read the error and try again. Which trade suits your agent is something only a measurement can settle, so the check keeps a log. For every call, code labels what the tool returned as data, nothing, or an error, so no model’s judgment is involved. Run the old version and the new one over the same questions and compare, per tool, how often calls fail and how often a failed call was followed by one that worked. The log also explains an answer that got thinner. If a new version’s answers say less, the log says why: if its calls failed more often, it had less to write from, and if they failed no more often, it wrote less of what it had.

Did the evidence arrive

Fig. 22: The answers match word for word, but the call log shows the communications lookup empty before the change and failing after, and across runs only its errors climb.

One step up is guardrail collision. Each guardrail is its own call to a smaller model, with a prompt of its own and a different incentive from the main agent’s: where the main agent looks for the highest leverage insight, the guardrail only decides whether the question or the answer breaks one of the rules configured for that agent, and returns pass, rewrite or block. Our eval runs log which rules fired, along with the four things that could have made them fire: the question and its answer, the model, the code, and the wording of the rules themselves. Hold three of them constant and change the fourth, and you can count the rules the change started hitting and the ones it stopped hitting. A guardrail is a probabilistic model too, like the judge, but its question is far narrower. Whether an answer breaks a rule can be decided from the answer alone, so the guardrail never needs the global corpus, which is what made the judge’s job impossible.

To better visualize it, think of the rules as a minefield you laid and each answer as a route across it. A few mines sit at the entrance and check the question before the agent says anything; most sit along the route and check the answer it produces. Swap the model or change the code, and the same question takes a new route, which may run into some mines more often and others less. Say your agent keeps promising customers release dates and setting off the no-roadmap-commitments rule, and your coding agent suggests giving it the product telemetry, so it can talk about what the customer uses today instead of speculating about what is coming. Then it starts setting off the telemetry-disclosure rule, and you loop to another change before this one ships. You can also move the mines themselves, by rewording the prose behind a guardrail’s policy, and observe the collision rate.

Guardrail collisions

Fig. 23: Above, a model or code change shifts the route from no-roadmap-commitments to telemetry-disclosure, and below, a reworded rule grows wider and wrongly stops an answer labeled should pass.

The rewrite rules lead to the next check. When the agent sets off telemetry-disclosure, the rule marks the passages that break it, and the whole answer goes to a larger model with those marks and one instruction: fix the marked passages and leave everything else alone. How well it follows that instruction is the rewriting model’s own behavior, and it can go wrong in two ways that both still read like a finished answer. Say the answer has five sections and quotes two raw usage figures, a monthly active user count and a count of unclaimed workspaces. The rewrite might reword the first figure and miss the second, two sections down. Or it might drop both figures along with three of the five sections, and open with “I removed the usage figures and kept the rest.” Call the first residue, where the problem the rule caught is still there, and the second collateral, where the problem is gone and so is much of the answer.

Code can find both failures, as long as each mark copies the exact words it objects to and is kept with the answer. The rewrite can then be checked against each half of what it was told. To check the fix, search the rewrite for each mark’s words: if they survive, or a number or date from them turns up in what replaced them, that is residue. To check the rest, compare the two answers sentence by sentence: every sentence no mark touched should come back unchanged, and any that is missing, reworded or new is collateral. A sentence with a mark in it may change, but it should not vanish. Replay the same saved answers and marks through two rewriting models, or two wordings of the instruction, and each test gives a rate you can compare.

Verify the repair

Fig. 24: Four possible rewrites of one marked answer meet two code checks, and residue fails only the mark search while collateral and a quiet loss fail only the sentence comparison.

Fact overlap is the widest of the four, and the most useful for charting performance against cost. It asks which pieces of the corpus the agent reached and relayed in its final answer. Every claim in the answer is traced to the item the tools returned that supports it: a field on the account record, a line in a call transcript, a comment on a ticket. Figures and dates match by value, names by entity, and a sentence by the passage it came from; when the agent cites the item it used, code only has to confirm the item says it. Each finding is keyed by its source, so two versions that word the same passage differently share one finding. Run several models side by side on the same code, or several commits on the same model, over one bank of cases. Nobody can compute which pieces, out of everything the agent could have read, belonged in the highest leverage subset, so each case gets a stand-in: the pool of every finding any version traced. The versions write the answer key for one another. Each version gets its share of the pool, along with what it missed and what only it found. Set the average share beside the cost per case, and you can read how much of the pool a cheaper configuration keeps for how much of the cost.

The versions weight the key as well. A share that counts every finding the same rewards padding, and it buries the one finding only the strongest model reached, which may be the one that justifies paying for it. So each finding carries a weight read off the answers: of the versions that reached it, how many led with it, in the opening or in the reason for the recommended move, and how many left it in the supporting detail. A finding only the strongest model reached, and led with, weighs as much as one every version led with, so its rarity does not count against it. A stray headcount that every version mentions and none leads with weighs little, so padding an answer barely moves its share. Two cautions remain. A finding no version reached never enters the pool, so the pool is a floor on what was there to find, and every share read off it runs high, though the floor rises with every version you run. And one model run twice may not return the same findings, so read each version against its own rerun first. Until a gap between two versions is larger than the gap between one version and itself, it may only be luck in which tool calls each run made.

Fact overlap

Fig. 25: Four versions are read against the pool of facts any of them found, and the cheaper model keeps a smaller share of it for a fraction of the cost.

You might wonder why anyone would try some of the changes above, when a great engineer embedded in the codebase could tell at a glance that they would make the agent worse. A coding agent often cannot tell, and with these checks it does not have to. It can suggest a change, run it over the bank of cases, read what the four checks say, keep it or revert it, and start again, for hours at a time. Claude Code or Codex’s /loop and /goal commands run exactly that kind of cycle, repeating one instruction on an interval or at its own pace until you stop it. Massaged this way, a non-verifiable domain becomes verifiable enough in aggregate to /loop or /goal on.

6. Train the Birds, Don’t Catch the Fish: Self-Healing Agents

Unfortunately for anyone who builds agents, nobody is ever truly satisfied with one. An artist eventually hands off a painting for others to admire, and an architect gets to walk through the building they conceived, but an agent never reaches a finished state you can hand off. You could say this has always been true of software, but I would argue it is more true of agents. Deterministic, UI-based software had more of a natural end state: you conceived an experience, made it beautiful and shipped it. An agent’s interface barely changes, since it is mostly just text, but the question behind it never goes away: could it do more for me, and what is it not surfacing? Born six hundred years too late for the Renaissance, you still want a polished artifact you can hand off, and there may be one, or at least a mirage of one: recursive self-improvement. Can you build an agent that keeps improving itself without anyone stepping in?

Perhaps someday, but that is a taller peak further on in the range, and you only need to get to the other side. A more measured mountain you can cross today is self-healing, an agent that inches itself back to a baseline you set from its observed performance. You get there by manually steering through every primitive you have (the model, tools, data sources, memory, prose, guardrails, execution horizon, token budget, and whatever comes out next). Every abstraction layer you build and every hill you climb gives you a chance to set that baseline, an expected score on a set of verifiable test cases. With a baseline, every new model becomes a trade you can read. You can see exactly where a cheaper, faster model or a newer, stronger one would land you on the cost and performance curve. The same goes for code changes to the agent harness. Your coding agent can try far more changes and keep only the ones that beat the baseline. The baseline itself gets more trustworthy, and harder to beat, as your bank of cases grows and the checks cover more of each answer.

Given this, a narrower change that could be automated is healing the prose, the way dreaming heals the memory. Today the prose surfaces are tuned by hand, and a new probabilistic model interprets each of them its own way, so changing the model means retuning every agent on the substrate by hand. An update harness could do that retuning instead. It would read the agent’s runs on the new model, notice where they fall short of the old baseline, and adjust the soft prose surfaces, the prompts, skills and tool descriptions at each layer, until they pass more often. Each change goes live only when it scores better than the last, and the harness keeps looping until the agent is back at the baseline or the changes have fallen below a diminishing-returns threshold you set.

That loop is what would give enterprises true model autonomy. No longer held hostage by the model their prompts were tuned for, they would know each model’s bang for the buck applied to their use case, and have an automated mechanism to close the gap to their baseline as far as another model allows.

To leave you with one last analogy, I recently had the pleasure of seeing Ukai, the ancient practice of fishing with birds, on the river at Arashiyama in Kyoto. In summer the fishermen take trained cormorants out after dark to catch fish for them, about five birds to a fisherman, each on its own cord. What stood out to me was the navigation. At first the fishermen painstakingly steered the boat up and down the river looking for a dense pocket of fish to show the birds, and settled on one once enough birds snapped at the fish in the water. Eventually it was the birds doing the navigation, steering the boat to one adjacent pocket after another and taking the fishermen to the fish. The whole river was too much for the birds to steer at the start, but given enough direction, they took over. Building great agents feels the same way: steer them until they steer themselves.

One night of ukai

Fig. 26: Early in the night the boat searches the river with the cormorants trailing on their cords, and later the birds find the fish first and the boat follows.

Appendix

SKILL.md, one agent’s manifest and system prompt.

---
name: deal-coach
model: <medium-frontier-model>
max_iterations: 30
channels: [slack, cli]

acl_enforcement: true
family: coach                    # inherit a shared grounding discipline
scope_bound_retrieval: strict    # only this caller's accounts
field_policy: commercial_only    # default-deny field allowlist

tools:
  exclude: [search_github_repos, query_telemetry_dashboard]
---

## Audience and scope

You coach the seller who invoked you, on accounts in their own book. They
are mid deal and short on time, and the question they ask is rarely the
whole question.

## The pipeline

Discovery, qualification, solution, proposal, close. Each stage has a job
to finish before the next can start, and most stalled deals stalled at an
earlier stage. A pricing fight is usually a discovery gap.

## Hard rules

1. Never narrate your own machinery.
2. Never invent a quote, a figure or a name. Every quote links to its source.
3. Say thin evidence out loud.
4. Never recommend a discount as the move to close. Go back to value.
5. Describe usage as a trend, never as raw figures.

## How to coach

1. Find where the deal really is. When the CRM and the last calls
   disagree, trust the calls.
2. Diagnose the kind of question: risk, stakeholders, competition, price,
   or expansion and renewal.
3. Recommend one move for this week, with the reason a senior seller
   would give.

## How to gather data

Memory first, then the CRM, then calls and emails, and usage only if the
question turns on adoption. Stop when you can answer.

## Output shape

What is true now, what it implies, and one next move, readable between
two meetings.

## Out of scope

If the account is not in their book, say so and name its owner. If the
question is not about moving a deal, point them to the right team.

GUARDRAIL.md, a shared rule that names the agents it applies to and runs on its own smaller, cheaper model.

---
id: no-discounting-recommendations
version: 2
title: Block ONLY when the agent RECOMMENDS a discount or price concession
       as the closing move. Pass factual pricing history and
       value articulation coaching.
stage: output
verdict_type: block
model: <small-fast-model>
applies_to: [deal-coach, deal-coach-pro, rollout-coach]
failure_policy: closed
---

Your ONLY job is to detect when the response recommends, as a sales
action, that the seller offer a discount, price drop, or pricing
concession as a path to close a deal.

DEFAULT TO PASS. Historical pricing facts, the customer's stated price
objection and escalation to deal desk are all legitimate. The narrow
violation is prescribing a price action.

reports/morning-brief.md, a skill with a schedule and a destination such as Slack, email, Confluence or Notion

---
name: morning-brief
cron: "0 8 * * 1-5"
outputs:
  - type: slack
    channel: "<channel-id>"
---

What happened across the GTM organization in the last 24 hours? Lead with
the most important signal, not a summary of everything.

tools.py, a sample custom tool for deal-coach that queries an embedded index of deals. This would be available to the agent in addition to the global tool set.

def search_similar_deals(situation: str, outcome: str = "any") -> str:
    """Closed deals most like the seller's situation, and how they ended."""
    hits = get_index("closed-deals").query(
        vector=embed(situation),
        top_k=5,
        include_metadata=True,
        filter=None if outcome == "any" else {"outcome": outcome},
    )
    return json.dumps([
        {**m["metadata"], "score": round(m["score"], 2)}   # summary, outcome, deciding_factor
        for m in hits["matches"]
    ])

TOOLS = [{
    "name": "search_similar_deals",
    "description": (
        "Find closed deals like the seller's situation and how each one ended, "
        "with the factor that decided it. Use to back a next move with what "
        "worked before. Returns summaries, not account names."
    ),
    "input_schema": {
        "type": "object",
        "properties": {
            "situation": {"type": "string"},
            "outcome": {"type": "string", "enum": ["won", "lost", "any"]},
        },
        "required": ["situation"],
    },
}]

HANDLERS = {"search_similar_deals": search_similar_deals}

The post Steer Your Agents Until They Steer Themselves appeared first on Postman Blog.

Read the whole story
alvinashcraft
5 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Build a Secure TypeScript MCP Client with Cross App Access (XAA)

1 Share

Enterprise apps rarely work alone. Imagine this scenario: HR wants to ensure that a new employee completes their new hire onboarding, and they want this verification in the company’s HR app so it’s tied to the employee’s personnel file. But the onboarding tasks come from a task-tracking tool with requirements defined by each department, from a different vendor and on a different domain than the HR app. It’s on IT to bridge the two.

Wiring the two together normally costs the user an OAuth consent screen or costs IT a shared service account. Cross App Access (XAA) removes that step. The company’s identity provider vouches for the user across the boundary, under a policy the admin sets in advance. In this tutorial, you’ll build that flow end-to-end in TypeScript.

You’ll build a Model Context Protocol (MCP) requesting app: a web app that signs a user in through an enterprise Identity Provider (IdP), exchanges that identity for a delegation token called an Identity Assertion JWT Authorization Grant (ID-JAG), trades the ID-JAG for an access token, and calls a protected MCP server to fetch real data. You’ll use the MCP TypeScript SDK, which added first-class Cross App Access support, and you’ll test everything against xaa.dev, the free XAA playground from the Okta Dev Advocacy team.

By the end of this post, you can:

  • Explain what XAA is and why AI agents and enterprise apps need it
  • Read an ID-JAG and understand key claims in it
  • Understand each step of the XAA flow through the SDK functions
  • Let the SDK’s CrossAppAccessProvider run the whole flow for you

Prerequisites

The code in this tutorial runs against @modelcontextprotocol/client version 2.0.0-alpha.2, openid-client 6, Express 4, and the xaa.dev playground services. The MCP SDK version is a prerelease of the v2 line, and APIs can change between prereleases, so check the sample repository for the exact dependency versions.

Table of Contents

What you’ll build: a TypeScript MCP requesting app

The sample app is an employee onboarding dashboard. After a single sign-in at the playground IdP, it fetches the user’s onboarding checklist from a protected MCP server and renders progress stats, an “up next” card, and the checklist itself. A “behind the scenes” panel shows all four XAA steps executing live, with timings and expandable decoded tokens, because watching the tokens flow is the best way to learn the protocol.

One naming note before the build: in this post, the MCP client and the requesting app are the same thing. The app is an MCP client toward the server it calls, and a requesting app in XAA terms, because it requests delegated access on behalf of the signed-in user.

The app spans four files:

  • src/server.ts: an Express server that handles the OpenID Connect sign-in with openid-client and streams the flow steps to the browser over server-sent events (SSE)
  • src/xaa.ts: the XAA flow via CrossAppAccessProvider, built on the MCP TypeScript SDK
  • src/config.ts: environment variables and app configuration
  • public/index.html: the dashboard

The complete project is on GitHub in the okta-xaa-typescript-mcp-sdk-example repository. Get it running first, then this post walks back through what each token in the chain does and how the SDK produces it.

Register your requesting app on xaa.dev

Rather than stand up an enterprise IdP and a protected resource yourself, you’ll code against xaa.dev, the free XAA playground from the Okta Dev Advocacy team. It hosts the three services you don’t have to build: the IdP (IdenX at https://idp.xaa.dev), the resource’s authorization server (https://auth.resource.xaa.dev), and a protected todo resource available as both a REST API and an MCP server (https://mcp.xaa.dev/mcp). You register your app once and get the credentials your project needs because XAA spans two trust domains: your app needs one identity at the IdP to sign users in and another at the resource’s authorization server to get access tokens. For a tour of the playground itself, read Introducing xaa.dev: A Playground for Cross App Access.

  1. Go to the requesting app registration page
  2. Enter your email address. It scopes which registered apps are visible to you; xaa.dev creates no account and sends no email.
  3. Select + Register New App and fill in the form:
    • Application Name: any label, for example, Onboarding App - Local Dev
    • Redirect URIs: http://localhost:3001/callback (the match is exact, including scheme, host, port, and path)
    • Connect to Resource: select the Todo MCP server resource (todo0-mcp) and keep the todos.read and mcp.access scopes
  4. Save the credentials from the confirmation modal

Registration gives you at least two sets of credentials, and mixing them up is the most common XAA mistake:

Client Credentials Identifies your app at
Main client client_id / client_secret the IdP, for sign-in and for the token exchange
Resource client resource_client_id / resource_client_secret (the ID looks like client_xxx-at-todo0-mcp) the resource’s authorization server

Note: Why separate credentials? The IdP and the resource’s authorization server are separate trust domains, so your app holds a separate identity at each. Using the main client’s credentials at the resource’s authorization server results in an invalid_client error.

Note: Some registrations also show a distinct pair for the token exchange itself. If yours does, keep it aside; the sample has a slot for it, and reuses the main client when you leave that slot empty.

Set up the TypeScript project

Clone the sample and install the dependencies:

git clone https://github.com/oktadev/okta-xaa-typescript-mcp-sdk-example.git
cd okta-xaa-typescript-mcp-sdk-example
npm install
cp .env.example .env

Fill in .env with the credentials from your registration:

# Main client: identifies your app at the IdP
XAA_CLIENT_ID=YOUR_CLIENT_ID
XAA_CLIENT_SECRET=YOUR_CLIENT_SECRET

# Token-exchange client, if your registration shows a separate pair for it.
# Leave both blank to reuse the main client above.
EXCHANGE_CLIENT_ID=
EXCHANGE_CLIENT_SECRET=

# Resource client: identifies your app at the resource's authorization server
MCP_CLIENT_ID=YOUR_RESOURCE_CLIENT_ID
MCP_CLIENT_SECRET=YOUR_RESOURCE_CLIENT_SECRET

# xaa.dev endpoints (no trailing slashes)
IDP_BASE_URL=https://idp.xaa.dev
AUTH_SERVER_URL=https://auth.resource.xaa.dev
MCP_SERVER_URL=https://mcp.xaa.dev/mcp

PORT=3001
BASE_URL=http://localhost:3001
XAA_SCOPE=todos.read mcp.access

Note: Keep the URLs free of trailing slashes. The IdP compares the audience value as an exact string, so https://auth.resource.xaa.dev/ with a slash fails where https://auth.resource.xaa.dev succeeds.

Note: The sample’s .gitignore excludes .env. Never commit client secrets, ID tokens, ID-JAGs, or access tokens to version control.

Run your TypeScript MCP app with xaa.dev

Start the server and open the dashboard:

npm start

Go to http://localhost:3001 and select Sign in with company SSO.

Employee Onboarding app welcome screen secured by Cross App Access, with a Sign in with company SSO button and enterprise identity provided by IdenX

IdenX accepts any email address, so no real credentials are involved. It then shows a Verify Your Identity screen asking for a verification code. The playground runs in demo mode and sends no email, so enter any six digits.

IdenX Verify Your Identity screen with a six-digit verification code field and a demo mode notice explaining that no email is sent and any six digits work

After sign-in, the app runs the flow automatically: four steps light up in order in the “behind the scenes” panel with real timings, and the onboarding checklist renders as soon as the last one delivers the data.

Employee onboarding dashboard showing a completed four-step Cross App Access flow with per-step timings, the issued access token, decoded token claims, and a to-do checklist fetched from the MCP server

You never saw a consent screen, and your app never held a credential for the todo service. That’s the point of XAA, and the panel shows how it happened.

The four-step XAA flow

The panel’s four cards are the whole protocol. Steps 1, 3, and 4 are standard OAuth 2.0; step 2 is the exchange XAA adds.

Step 1: User signs in (OpenID Connect + PKCE)
  Browser -> IdP                            -> your app gets an ID token

Step 2: Token exchange (RFC 8693)
  Your app -> IdP token endpoint            -> ID-JAG

Step 3: JWT bearer grant (RFC 7523)
  Your app -> resource authorization server -> access token

Step 4: Call the resource (RFC 6750)
  Your app -> MCP server (Bearer token)     -> data

The user authenticates once with the IdP, and your app receives an ID token. Your app exchanges that ID token at the IdP for an ID-JAG using OAuth 2.0 Token Exchange (Request for Comments (RFC) 8693). Your app presents the ID-JAG to the resource’s authorization server using the JWT bearer grant (RFC 7523) and receives a scoped access token. Finally, your app calls the protected resource with that token as a standard Bearer credential (RFC 6750).

The Cross App Access flow in the TypeScript MCP app, in which the user, web app, IdP, authorization server, and MCP server exchange an ID token, an ID-JAG, and a Bearer access token across four steps

Inspect the tokens in the dashboard

Select any step card to expand its decoded token. The next section walks through what the ID-JAG’s claims mean; for now, the one thing worth confirming is that the access token’s aud matches the ID-JAG’s resource, byte for byte. That match is the delegation chain holding together, and a mismatch is the most common cause of a 401 from the resource server.

Below the flow steps, the ACCESS TOKEN card displays the Bearer token issued in step 3, with an expand toggle to inspect the full JWT. The TOKEN CLAIMS card decodes the same token and surfaces iss, aud, sub, and scope. Together, they confirm the right issuer signed the token, it targets the MCP server, it carries the user’s identity, and the authorization server granted the scopes you asked for.

Select 🔄 Re-run (SDK discovers the auth server) at any time to replay the flow and watch CrossAppAccessProvider handle discovery automatically.

Note: The dashboard’s token inspector exists for learning and local debugging. Keep raw tokens out of production interfaces and logs; when troubleshooting in production, log redacted identifiers and non-sensitive claims instead.

Note: This app is read-only by design. It requests todos.read and mcp.access, and nothing else. Least privilege applies to AI agents and requesting apps the same way it applies to users.

ID-JAG: the token that carries identity across apps

The ID-JAG is a new token type introduced by XAA. It’s a signed JSON Web Token (JWT) that the IdP issues when your app exchanges the user’s ID token for it. Think of it as a sealed envelope from the IdP that says: “This app acts for this user, toward this specific resource, with these scopes, for the next five minutes.”

Here is a decoded ID-JAG from the app you just ran:

{
  "iss": "https://idp.xaa.dev",
  "sub": "usr_1a2b3c4d5e6f",
  "aud": "https://auth.resource.xaa.dev",
  "resource": "https://mcp.xaa.dev/mcp",
  "client_id": "client_1a2b3c4d-at-todo0-mcp",
  "scope": "todos.read mcp.access",
  "email": "user@example.com",
  "jti": "69586a9a-b962-4f5d-971f-40f12173bcf2",
  "iat": 1783333416,
  "exp": 1783333716
}

Most of those claims are standard JWT fields. Three decide whether your integration works:

  • aud names the token’s consumer: the authorization server that validates this ID-JAG. It isn’t the resource itself.
  • resource names the target API. The authorization server copies this value into the access token’s aud claim, so whatever you put here is what the resource validates later.
  • client_id is your app’s identity at the resource’s authorization server, not at the IdP. That’s the two-client model from registration, showing up in the token.

Get aud and resource backwards and the exchange either fails outright or succeeds into a token the resource server rejects, which is the error that costs the most debugging time.

The JWT header also carries "typ": "oauth-id-jag+jwt", and authorization servers reject anything else. That prevents attackers from replaying other JWTs, such as ID tokens, as authorization grants. The five-minute exp window and the one-time jti close off replay from the other direction.

Why Cross App Access matters for AI agents

You just watched one user action cross one app boundary. Enterprise software rarely stops at one. A single action fans out across many systems, and AI agents make the fan-out constant: an assistant reads a document, files a ticket, checks a calendar, and logs an audit event, all on behalf of one person.

Each of those hops needs the user’s identity. The traditional answer is a consent screen at every boundary. Users click through prompts they don’t read, and IT teams lose visibility into which app talks to which other app, because the authorization rests on personal consent. Some teams give up and share a service account, which puts a single overprivileged credential in front of everyone’s data.

XAA replaces that consent step with a policy the admin configures in advance. Your enterprise IdP already knows who the user is, because the user signed in this morning, so it vouches for that identity across the boundary with a short-lived signed token. The admin decides which app connects to which resource, with which scopes. The user signs in once, and every hop after that is an auditable cryptographic handoff.

Cross App Access is the industry term for a pattern built on the Identity Assertion Authorization Grant specification, an active Internet-Draft at the Internet Engineering Task Force (IETF) that defines the ID-JAG token and the exchange flow that XAA uses. To go deeper on the protocol itself, read Build Secure Agent-to-App Connections with Cross App Access (XAA).

The Model Context Protocol and its TypeScript SDK

Before diving into code, a quick word on the other protocol in this tutorial. The Model Context Protocol is an open standard created by Anthropic that continues to evolve, providing AI applications with a common way to connect to tools and data. An MCP server exposes capabilities (tools to call, resources to read, prompts to use), and an MCP client connects to those servers over a standard transport. Instead of building one custom integration per data source, build to one protocol so any MCP-capable AI application can use it. The MCP TypeScript SDK is the official implementation for JavaScript and TypeScript developers, handling protocol details such as transports, the initialization handshake, message schemas, and authorization.

The MCP community evolves the specification through Specification Enhancement Proposals (SEPs). Cross App Access support entered the protocol as Specification Enhancement Proposal (SEP) 990, Enterprise Managed Authorization, and the TypeScript SDK ships the implementation in its crossAppAccess module.

With SEP-990 implemented in the SDK, an XAA-enabled MCP client differs from a plain one by a single authProvider option and one callback, plus any compatibility adjustments your authorization server needs. The protocol work (discovery, token exchange, the JWT bearer grant, and retries) lives in the SDK rather than in your app.

Inside the MCP TS SDK’s crossAppAccess module

Three pieces of the TypeScript SDK matter for this tutorial:

  • discoverAndRequestJwtAuthGrant() performs step 2. It discovers the IdP’s token endpoint from its metadata, then sends an RFC 8693 token-exchange request and returns the ID-JAG. A sibling function, requestJwtAuthorizationGrant(), skips discovery when you already know the token endpoint.
  • exchangeJwtAuthGrant() performs step 3: given the resource authorization server’s token endpoint, it presents the ID-JAG with the RFC 7523 JWT bearer grant and returns the access token. Unlike the step 2 function, it does no discovery of its own.
  • CrossAppAccessProvider is the production path. It plugs into the SDK’s transport as an OAuthClientProvider and runs steps 2 through 4 automatically: it calls the MCP server, receives a 401 challenge, discovers the authorization server through protected resource metadata (RFC 9728), invokes your callback to obtain a fresh ID-JAG, exchanges it, and retries the request with the new access token.

You see discoverAndRequestJwtAuthGrant() in step 2 and the provider in the final section. exchangeJwtAuthGrant() is the standalone way to do step 3; the sample never calls it, because the provider builds that same JWT bearer request itself.

Walk through the XAA flow in TypeScript

The next four sections take each step in turn, so that you can trace every token in the chain. The fifth shows CrossAppAccessProvider, which drives steps 2 through 4 from one configuration. The code below is adapted from src/server.ts and src/xaa.ts, trimmed to the lines that matter for each step.

Sign the user in with OpenID Connect and PKCE

Step 1 is a standard OpenID Connect Authorization Code flow with PKCE, and standard means you shouldn’t write it yourself. The sample uses openid-client, which handles discovery, the PKCE challenge, and the code exchange. The piece XAA cares about is the ID token in the response, because it becomes the input to step 2.

Discovery returns a configuration object that the other calls take as their first argument. In src/server.ts:

import * as client from 'openid-client';

const config = await client.discovery(
  new URL('https://idp.xaa.dev'),
  process.env.XAA_CLIENT_ID,
  process.env.XAA_CLIENT_SECRET,
);

The /login route generates a fresh PKCE verifier, state, and nonce, stores them in the session, and builds the authorization URL:

const verifier = client.randomPKCECodeVerifier();
session.pkceVerifier = verifier;
session.state = client.randomState();
session.nonce = client.randomNonce();

const url = client.buildAuthorizationUrl(config, {
  redirect_uri: REDIRECT_URI,
  scope: 'openid email profile',
  state: session.state,
  nonce: session.nonce,
  code_challenge: await client.calculatePKCECodeChallenge(verifier),
  code_challenge_method: 'S256',
});

res.redirect(url.href);

The /callback route then exchanges the code for tokens in one call:

const tokens = await client.authorizationCodeGrant(
  config,
  new URL(req.originalUrl, BASE_URL),
  {
    pkceCodeVerifier: session.pkceVerifier,
    expectedState: session.state,
    expectedNonce: session.nonce,
  },
);

session.idToken = tokens.id_token;
session.claims = tokens.claims();

That one call verifies the ID token’s signature against the IdP’s published keys, checks the iss, aud, and exp claims, and confirms the state and nonce match what you stored. Hand-rolled sign-in code tends to decode the ID token payload without verifying any of it, which means trusting claims from a token you haven’t proven came from your IdP. Here, the ID token is the subject_token for step 2 and the root of the whole delegation chain, so that verification is load-bearing.

The app keeps the ID token in a server-side session. It never sends the ID token to the resource server; the token’s only job now is to prove the user’s identity to the IdP during token exchange.

Production note: Configure session cookies with HttpOnly, Secure, and an appropriate SameSite value, and keep client secrets out of browser code. The sample stores sessions in memory, which is fine for a local demo and wrong for production: use a session store so tokens survive a restart and scale past one process.

Exchange the ID token for an ID-JAG

Step 2 is the XAA step. One SDK call performs the whole RFC 8693 token exchange:

import { discoverAndRequestJwtAuthGrant } from '@modelcontextprotocol/client';

const jag = await discoverAndRequestJwtAuthGrant({
  idpUrl: 'https://idp.xaa.dev',
  audience: 'https://auth.resource.xaa.dev', // who validates the ID-JAG
  resource: 'https://mcp.xaa.dev/mcp',       // what the access token targets
  idToken: session.idToken,
  clientId: process.env.XAA_CLIENT_ID,       // main client
  clientSecret: process.env.XAA_CLIENT_SECRET,
  scope: 'todos.read mcp.access',
});

console.log(jag.jwtAuthGrant); // the ID-JAG (a JWT)
console.log(jag.expiresIn);    // 300 seconds

Under the hood, the function discovers the IdP’s token endpoint from its metadata, then sends a form-encoded POST with grant_type=urn:ietf:params:oauth:grant-type:token-exchange, your ID token as the subject_token, and requested_token_type=urn:ietf:params:oauth:token-type:id-jag. The response’s access_token field carries the ID-JAG. The naming feels odd, but RFC 8693 reuses the standard token response shape, and the token_type value N_A signals that this token isn’t a bearer credential.

Pay attention to audience versus resource, because they answer different questions. audience names who validates this ID-JAG in step 3: the authorization server. resource names what your access token targets in step 4: the MCP server. The authorization server copies the resource verbatim into the access token’s aud claim, so an incorrect resource here later surfaces as a confusing 401 from the resource server.

One optimization worth noting: the discovery call incurs an extra network round trip on every run. If your server already fetched the IdP metadata (this app caches it during sign-in), call requestJwtAuthorizationGrant() with the known tokenEndpoint instead and skip rediscovery. Caching static configuration is fine; the tokens themselves are always requested live.

Exchange the ID-JAG for an access token

In step 3, your app presents the ID-JAG to the resource’s authorization server with the RFC 7523 JWT bearer grant, authenticating with the resource client credentials. The request posts grant_type=urn:ietf:params:oauth:grant-type:jwt-bearer with the ID-JAG as the assertion, and the response is an ordinary OAuth token response: an access_token, token_type of Bearer, and an expires_in lifetime worth reading rather than assuming.

You don’t write that request yourself. CrossAppAccessProvider makes it, and the last section of this walkthrough shows how to configure it. Two details about this step save you real debugging time:

  1. Client authentication method. Developer-registered clients on xaa.dev use client_secret_post, which means credentials belong in the request body. The provider declares client_secret_basic (an Authorization: Basic header) by default, so the client information has to name the method explicitly.
  2. Send the scope parameter. The provider only sends a scope when the authorization server’s metadata advertises one, and xaa.dev’s does not. Omit it and the authorization server issues an access token with an empty scope, after which the failure is quiet. The MCP server still completes the handshake, resources/list still returns the resource names, and reading the todos still returns HTTP 200 with a JSON-RPC result. The rejection hides inside the resource payload: {"error":"Unauthorized","message":"Invalid or expired token"}. Nothing throws, so your app parses that error object instead of a todo list and renders an empty checklist. Request the scopes you need, then verify the scope claim in the decoded access token.

The resulting access token is itself a JWT. Its aud claim matches the resource you sent in step 2, its sub identifies the user, and its client_id is the resource client. This is the delegation chain made visible: user identity from step 1, admin policy from the resource connection, and app identity from your registration, all cryptographically bound into one credential.

Fetch data from the MCP server

With a token in hand, step 4 is regular MCP SDK code. Given a connected Client, which the next section builds, reading a resource takes one call:

const read = await client.readResource({ uri: 'todo0://todos' });

// A resource's contents can be text or binary, so narrow before parsing.
const first = read.contents[0];
const todos = first && 'text' in first ? JSON.parse(first.text) : [];

readResource() returns the user’s todo list as JSON. The playground’s MCP server exposes the todos as an MCP resource (read-only data at a URI) rather than a tool. Both primitives ride the same authenticated pipeline, so switching to a tool call is a one-line change to client.callTool() when your resource server offers tools.

Notice that the access token never appears in this code. You can attach it to the transport yourself through requestInit.headers, but the sample gives the transport an authProvider instead, and the SDK fetches the token, attaches it, and refreshes it when it expires.

CrossAppAccessProvider runs the whole flow

The sections above describe what each step sends and receives. Here is the code that runs all three, from src/xaa.ts:

import {
  Client,
  CrossAppAccessProvider,
  StreamableHTTPClientTransport,
  discoverAndRequestJwtAuthGrant,
  requestJwtAuthorizationGrant,
} from '@modelcontextprotocol/client';

const provider = new CrossAppAccessProvider({
  // Called when the transport needs an ID-JAG. The provider has already
  // discovered the authorization server and resource via RFC 9728.
  assertion: async (ctx) => {
    // idpTokenEndpoint is the IdP token endpoint this app cached during
    // sign-in, passed in by the caller. Falls back to discovery without it.
    const jagOptions = {
      audience: ctx.authorizationServerUrl.replace(/\/+$/, ''),
      resource: ctx.resourceUrl,
      idToken: await getIdTokenFromSession(),
      clientId: process.env.XAA_CLIENT_ID,
      clientSecret: process.env.XAA_CLIENT_SECRET,
      scope: ctx.scope ?? 'todos.read mcp.access',
      fetchFn: ctx.fetchFn,
    };
    const jag = idpTokenEndpoint
      ? await requestJwtAuthorizationGrant({ ...jagOptions, tokenEndpoint: idpTokenEndpoint })
      : await discoverAndRequestJwtAuthGrant({ ...jagOptions, idpUrl: 'https://idp.xaa.dev' });
    return jag.jwtAuthGrant;
  },
  clientId: process.env.MCP_CLIENT_ID,
  clientSecret: process.env.MCP_CLIENT_SECRET,
});

// Two adjustments for xaa.dev developer clients. First, they authenticate
// with client_secret_post, while the provider declares client_secret_basic
// by default. Declaring the method on the client information makes the SDK
// select the right one.
provider.saveClientInformation({
  client_id: process.env.MCP_CLIENT_ID,
  client_secret: process.env.MCP_CLIENT_SECRET,
  token_endpoint_auth_method: 'client_secret_post',
} as Parameters<typeof provider.saveClientInformation>[0]);

// Second, the provider only sends a scope when the server's metadata
// advertises one, and xaa.dev's does not. Without a scope parameter, the
// authorization server issues an empty-scope token, so add it here.
const prepareTokenRequest = provider.prepareTokenRequest.bind(provider);
provider.prepareTokenRequest = async (scope) => {
  const params = await prepareTokenRequest(scope ?? 'todos.read mcp.access');
  if (!params.has('scope')) params.set('scope', 'todos.read mcp.access');
  return params;
};

const transport = new StreamableHTTPClientTransport(
  new URL('https://mcp.xaa.dev/mcp'),
  { authProvider: provider },
);

const client = new Client({ name: 'xaa-requesting-app-typescript', version: '1.0.0' });
await client.connect(transport); // 401 -> discovery -> ID-JAG -> token -> retry

Walk through what the provider does on that one connect() call. It sends the first request without a token and receives a 401 with a WWW-Authenticate header pointing at the server’s protected resource metadata. It fetches that metadata, learns which authorization server protects this resource, and then fetches the authorization server’s metadata to find the token endpoint. It calls your assertion callback with the discovered URLs, exchanges the returned ID-JAG for an access token, stores the token, and then retries the original request. When the access token later expires, the next 401 response triggers the same sequence again, without any code from you. Both adjustments in the snippet come from testing this flow against the live playground.

Nothing in this snippet names auth.resource.xaa.dev. The provider discovered it, so the same client code works against any spec-compliant protected MCP server.

Where XAA fits in your real architecture

The playground stands in for real systems, and the mapping is direct. IdenX stands in for your production IdP. Okta offers Cross App Access as an Early Access feature, and the token exchange in this tutorial works the same way against an Okta org once you enable the feature; check the Okta Cross App Access documentation for current availability. To try it against a real org, sign up for an Okta Integrator Free Plan. The todo MCP server serves as an enterprise resource you can make available to agents and apps. Your requesting app serves as the AI agent or SaaS integration acting on the user’s behalf.

The read-only scope in this tutorial is a deliberate starting point, not a limitation of the protocol. Scopes are strings that your resource server defines and an admin grants. When you’re ready to build the other side of the boundary, the resource app guides show how to validate ID-JAGs and issue access tokens from your own authorization server, whether your app federates with OpenID Connect (OIDC) or Security Assertion Markup Language (SAML).

Learn more about Cross App Access, ID-JAG, and MCP

The complete source for this tutorial is in the okta-xaa-typescript-mcp-sdk-example repository. To keep exploring:

Remember to follow us on LinkedIn, X, and subscribe to our YouTube channel for more exciting content. We also want to hear from you about the topics you’d like to see and any questions you may have. Leave us a comment below!

Read the whole story
alvinashcraft
5 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

An Anthropic AI model sent a false homicide tip to Philadelphia police

1 Share
Anthropic did not discover this behavior until over two months after its AI submitted the false tip.
Read the whole story
alvinashcraft
5 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

From H-1B to CEO: How Satya Nadella traveled the path the U.S. just cut off for Microsoft workers

1 Share
President Trump applauds the six recipients of this year’s National Medals of Science and National Medals of Technology and Innovation on Thursday in Washington, D.C. Five were born outside the U.S., including Microsoft’s Satya Nadella, far left. (White House Video Screenshot)

Satya Nadella arrived in the U.S. from India in 1988 to study computer science. Over the next three decades, he went from a temporary work visa to permanent residency to American citizenship, and rose up the ranks at Microsoft to become only the third CEO in its history.

“I don’t think my story would be possible anywhere else,” Nadella wrote in his 2017 book, “Hit Refresh.” “I am proud today to call myself an American citizen.”

On Thursday, President Trump personally handed the Microsoft CEO the National Medal of Technology and Innovation. The award cited his “transformative leadership of a critical American technology company, which led to its rejuvenation for a new era of cloud computing.”

Nadella thanked Trump on X, writing that he was “humbled to stand alongside so many giants of American innovation.” The other honorees were Dell Technologies CEO Michael Dell, who also received the technology medal; and Elon Musk, Google co-founder Sergey Brin, Nvidia CEO Jensen Huang and AMD CEO Lisa Su, who received the National Medal of Science.

Five of the six were born outside the country.

Hours earlier, the Trump administration had suspended Microsoft from the federal program it uses to help foreign workers take the same path to permanent residency. Vice President JD Vance said the company had abused the system by laying off American workers while hiring foreign workers, and called H-1B workers “foreign indentured servants.”

Vice President JD Vance speaks at a White House news conference on visa fraud Thursday as White House Deputy Chief of Staff Stephen Miller, center, and Attorney General Todd Blanche, right, look on. (White House Video Screenshot)

“Our message to Microsoft is you’re a great American company, but you’ve got to hire great American workers,” Vance said at a White House news conference.

Two different programs: The suspension doesn’t stop Microsoft from hiring foreign workers on H-1B visas. Those are temporary work visas, typically granted for three years at a time and generally limited to six years, with some exceptions. They’re overseen by U.S. Citizenship and Immigration Services.

Amazon is the bigger user of H-1B visas. It received more than 19,000 approvals for H-1B petitions in fiscal 2025, about three times Microsoft’s total, according to an analysis of USCIS data by the National Foundation for American Policy. The totals include renewals and extensions, so they don’t represent individual workers.

The program to which Microsoft lost access comes later in the process.

Known as PERM, it’s run by the Labor Department and is usually the first step when a company sponsors an employee to stay in the U.S. permanently. The employer has to advertise the job and show that no qualified, available U.S. worker can fill it. Most applications are for people already working in the U.S. on H-1B visas.

Microsoft is the biggest user of PERM. The Labor Department ruled on 3,161 of its applications in the fiscal year that ended in September 2025, more than for any other employer in the country, according to a GeekWire analysis of the agency’s data. Microsoft also led in the first nine months of fiscal 2026, with 1,683 applications decided through June.

Amazon and Google filed thousands of PERM applications in 2022 and early 2023, then largely stopped. The Labor Department ruled on 39 applications from Amazon and its subsidiaries last fiscal year, including Whole Foods, Audible and Zappos, and three from Google.

The suspension could hit hardest for Microsoft employees already on H-1B visas. Many count on the company to start the permanent residency process in time to extend their visas past the usual six-year limit, and the Labor Department says it won’t process Microsoft’s pending applications either.

Some Microsoft employees on H-1B visas were surprised by the news and unsure whether it would affect them, The Seattle Times reported. The paper noted that the White House hasn’t provided evidence of fraud by Microsoft.

Microsoft’s response: In a statement posted Thursday, Microsoft pushed back on criticism of its hiring and said it looks forward to “providing the Administration with additional information.” The statement focused on H-1B visas and didn’t address the PERM suspension specifically.

The company said about 80% of the roughly 6,000 H-1B applications it filed last fiscal year were to extend or change the status of people who already work at Microsoft. “These were not to hire new people,” it said.

The rest were for new hires already living in the U.S. legally, a group equal to about 1% of Microsoft’s U.S. workforce. “They are not new arrivals to our country,” the company said.

Microsoft also said it files H-1B petitions only for workers who meet “the rigorous standards of this visa category,” pays them “the same as any other employees doing comparable work,” and offers wages “among the highest of all H-1B filings.”

The vast majority of its U.S. employees are Americans, the company said.

Nadella’s own path: Nadella’s route to citizenship wasn’t a straight line, as he described in “Hit Refresh.” He was working at Microsoft when he married his wife, Anu, in India in December 1992. The next summer, her application for a visa to join him in the U.S. was rejected because she was married to a permanent resident.

Spouses of permanent residents faced long waits, and Microsoft’s immigration lawyer told Nadella it could take five years or more to bring her to the U.S.

Nadella considered quitting Microsoft and moving back to India. Then the lawyer suggested a different approach: give up his permanent residency and go back to a temporary H-1B visa, which would let his wife join him right away.

In June 1994, Nadella went to the U.S. embassy in New Delhi, past the long lines of people hoping to get visas, and told a clerk he wanted to give up his permanent residency. The clerk was dumbfounded, Nadella wrote. The plan worked, and Anu joined him in Seattle, where they started a family. Nadella eventually got back on the path to permanent residency and later became a U.S. citizen.

But in the meantime, back in Redmond, Nadella became known around the Microsoft campus as the guy who gave up his green card. Colleagues called him for advice. One, Kunal Bahl, later left Microsoft when his H-1B ran out before his permanent residency came through. He returned to India and founded the e-commerce company Snapdeal.

“Such is the perverse logic of this immigration law,” Nadella wrote.

Read the whole story
alvinashcraft
6 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Spaces, Tabs, Panes, and Precise Words: Why I Break Agent Work Out the Way I Do

1 Share

In the previous post I toured Herdr and how I lay it out: one space per folder, one tab per kind of work, agents in tall columns, shells in short stacks, all of it spread across a 5120×1440 ultrawide. That’s the what. This one is the why, and then the part that actually decides whether any of it pays off: how I talk to the agents once they’re in their panes.

Short version: the layout isn’t decoration. It’s context management for two kinds of brains, mine and the models’, and the prompts are the same idea at a smaller scale. Every tool in this stack, Claude Code, Codex, Cursor, and the human, gets worse the more unrelated junk is in its head. So I keep the junk out, structurally first and then verbally.

Part 1: Why I break it out like this

Spaces are folders because the folder is the agent’s whole world

An agent’s working directory isn’t a detail. It’s the edge of its universe. Whatever it can read, grep, build, and break starts at that folder. So when I make one space per folder, I’m not organizing windows. I’m drawing the same boundary the agent already lives inside, and making it visible to me.

That buys three things:

  1. The sidebar is a readout of my working memory. Ten spaces means ten things in flight. If a folder isn’t in the sidebar, I’m not working on it. When I’m done, the space goes away. There’s no “misc” space and no “scratch” space, because “misc” is where context goes to die.
  2. Blast radius is legible. If something is going sideways in ledger-service, the dot next to ledger-service says so. Nothing in that space can be quietly editing storefront-web, because it doesn’t live there.
  3. It pairs with worktrees. When I need two implementation agents on the same repo at once, each gets its own worktree and its own branch, and in Herdr that’s just another folder, so another space (herdr worktree create does this directly). I made the full case in Giving Every Agent Its Own Branch: Git Worktrees for Parallel AI Work: “Run two coding agents against a single working tree and they will fight.” Spaces-as-folders keeps the referee’s view honest.

Tabs are modes because each mode is a different verb

I name tabs domain-modeling, implementation, and pr-reviews because those are three different jobs. Each one comes with a different posture from me and a different register of prompt to the agent.

TabThe verbMy postureWhat the agent is allowed to do
domain-modelingthinkSkeptical, SocraticRead, reason, write a model. No code. Plan mode.
implementationchangeDirective, fencedEdit inside an explicit scope, after an approved plan
pr-reviewsjudgeAdversarial, evidentiaryRead the diff, report findings. No fixing.

The failure mode I’m designing against is verb drift. You start a conversation reviewing a PR, you say “huh, good catch,” and three messages later the “reviewer” is rewriting the module, now with the confidence of something that just graded its own homework. Separate tabs, and separate agent sessions in them, make the verb change explicit. If I want a fix, I go to the implementation tab and ask for one there, with a scope. It’s the same rule I put on the read-only reviewer agents in Ask, Plan, Confirm: Making Agents Stop Before They Start: the reviewer doesn’t get to “slide from ‘here is a finding’ into ‘and I fixed it’ without crossing the same line everyone else crosses.” The tab is that line, drawn on screen.

There’s a human benefit too. When I switch to a tab named pr-reviews, my brain switches hats before I’ve read a word. Mode labels are cheap and they work.

A pane is one agent, one context window, one job

Every agent pane gets one job for the life of that conversation. When the job is done, the conversation ends, and the next job gets a fresh agent. Context windows are big now, but they aren’t free and they aren’t clean. Every stale plan, abandoned approach, and “actually, ignore that” sits in there and pulls on the next answer. A pane per job is the cheapest context hygiene there is.

The panes next to an agent are there for verification, not decoration:

  • In domain-modeling, the glossary and domain notes sit beside the agent, so I can check that it’s using the words we agreed on.
  • In implementation, a diffstat and a shell sit beside the agents, so I can see what actually changed, not what the agent says changed.
  • In pr-reviews, the raw diff sits beside the reviewer, so every file:line it cites can be checked by moving my eyes and nothing else.

Trust, but keep the evidence in the same field of view.

Tall for agents, short for shells, and why 32:9 matters

Agents emit long, structured output (plans, findings, tables), so they get full-height vertical splits. Shells and watchers are glanceable, so they get short horizontal stacks in a utility column. That isn’t a style choice. It’s matching pane shape to how much reading each pane needs.

The ultrawide makes it work without compromise. At 5120×1440 I get three or four full-height agent columns and the utility column and the sidebar, with nothing hidden behind a zoom or another tab. That matters for one reason above all: a blocked agent is a cost that accrues silently. Herdr’s sidebar tells me that something is blocked. Having it physically on screen tells me why, in peripheral vision, without moving. The less effort it takes to notice, the shorter the stall.

Part 2: Keeping Claude, Codex, and Cursor effective and accurate

All of the above is plumbing. The water is language.

I’ve been on this soapbox a while. In Precision in Words, Precision in Code: The Power of Writing in Modern Development I argued that “Writing isn’t just ‘extra work’—it’s the work that clarifies, simplifies, and accelerates everything else,” and in AI Prompt Engineering: Mastering Language Constructs I went into how the actual constructs, imperatives, conditionals, and contextual markers, shape what comes back. With agents that act, not just answer, the stakes go up. A vague word in a chat gets you a vague paragraph. A vague word in an agent prompt gets you a nine-file diff.

Here’s what I actually do. Every example below is a real prompt and a real response, captured for the screenshots in this repo with Claude Code in plan mode against throwaway demo projects.

1. Use the domain’s words, and ban the overloaded ones

Accurate language starts with the same word meaning the same thing every time: in the docs, in the code, in the prompt, in the review. That’s just ubiquitous language from domain-driven design, and agents need it more than humans do, because they fill every gap with the most statistically likely meaning, which is rarely your meaning.

So I keep a glossary in the notes repo (more on that in Notes Repos as Shared Project Context, where the glossary is described as “cheap insurance against confident mistakes”), and I put the words in the prompt, including the ones that are off limits:

Read docs/domain.md and internal/ledger/entry.go. Using only the domain terms Account, Entry, Journal, and Posting (never ‘transaction’), describe a Posting aggregate: its invariants, what it owns, and what it must reject. Do not write code. Output a short markdown model.

“Transaction” is banned because in a ledger it’s hopelessly overloaded: a database transaction, a business event, a bank-statement line. Pick one meaning, give it its own word, and forbid the ambiguous one. Here’s what came back:

Two things to notice. The model sticks to Account, Entry, Journal, Posting the whole way through. And it marks which rules it inferred from general double-entry practice versus which came from the repo: “Rules marked (inferred) come from standard double-entry practice, not the repo.” That distinction is gold in domain modeling, and it showed up because the prompt was explicit about where the truth lives (those two files).

2. Lead with scope, as the first word

My implementation prompts literally start with the word Scope:.

Scope: cmd/main.go only. Plan a minimal CLI entrypoint that loads a Journal from a JSON file and prints the balance per Account. Do not touch internal/ledger. Keep the planned diff under 50 lines.

Scope first, because everything after it gets read in its light. Then a positive fence (only this file), a negative fence (not that package), and a size limit stated as a number, not a vibe. “Keep it small” means nothing. “Under 50 lines” is checkable. I went deeper on why scope and diff size are the main levers in Hurting or Helping Devs?.

The payoff:

The repo had no go.mod, so cmd/main.go couldn’t import internal/ledger. An unfenced agent “fixes” that by adding a module file, rewiring imports, and maybe stubbing a Journal type in the package I’d fenced off. This one stopped and asked, offered the in-scope option as the recommendation, and named the out-of-scope option as out of scope: “This touches a second file, which goes beyond the cmd/main.go-only scope.” Herdr flagged it blocked, and I could see it on the ultrawide without switching anything. That’s the whole system working: precise words create the stop, and the layout makes the stop cheap.

3. Pick verbs that mean exactly one thing

plan, implement, review, and fix are different verbs, and I never let one stand in for another.

  • Plan means produce a plan and change nothing. I back it with tooling: agents start in plan mode (claude --permission-mode plan; Codex and Cursor have their own read-only and ask-first modes), so the verb and the permission agree.
  • Implement means change things inside the approved plan’s scope, and only after an explicit yes. Approval is scoped to the plan that was approved. This is the three-beat gate from Ask, Plan, Confirm.
  • Review means read and report. It doesn’t mean fix.
  • Fix means a new implementation task, with its own scope, in the implementation tab.

When the verb is precise, the agent’s job is precise, and so is my evaluation of whether it did the job.

4. Specify the shape of the answer

If I can’t check an answer quickly, I’ll check it badly. So I specify the output format in a way that’s easy to verify against the pane next door:

Review the diff main…feature/posting-aggregate. Classify every change as additive, mutative, or destructive. Flag any naming that drifts from docs/domain.md. Report findings as a numbered list with file:line. Do not propose rewrites outside the diff.

“Additive, mutative, or destructive” isn’t decoration either. It’s the vocabulary from Additive vs. Mutative vs. Destructive Code Changes (and Why AI Agents Love the Wrong One at 2:13AM), and giving the reviewer that vocabulary makes it classify instead of just commenting.

And because the question was precise, the reviewer had attention left over for what mattered. Under “side notes” it flagged that any Direction other than the exact string "debit", including "" or "Debit", gets counted as a credit, and that an empty Posting passes validation as balanced. Neither is a naming issue. Both are exactly the kind of quiet behavioral drift that post warned about: code that looks right and behaves like an alien artifact. A vague “please review this PR” tends to bury that under eleven style nits.

The same pattern works for any review lens. On the storefront demo, “integer cents only, no floating point until display formatting. Numbered findings with file:line, severity first” produced a short, ranked list I could verify in about thirty seconds.

5. Front-load the context instead of looping for it

A thin prompt plus five rounds of correction is worse than one well-fed prompt, and more expensive. I made that argument at length in “Loop Engineering” Is Mostly Just Broken SDLC Wearing a Costume: “A well-fed single pass beats a starved five-pass loop most of the time.” In practice that means every prompt names where the truth lives: the files to read, the glossary, the plan in the notes repo, the decision that was already made and shouldn’t be re-litigated.

It also means the environment has to be ready. An agent that can’t build, test, or reach the systems it needs will improvise, and improvisation is the opposite of accuracy. The checklist I use for that is in Agent-Ready Local Environments and MCP: “An agent with broad access and no rules will eventually do something technically impressive and socially awful.”

6. Same template, any agent

Claude Code, Codex, and Cursor have different personalities and different defaults, but the prompt shape carries across them. Herdr doesn’t care which one is in a pane. It detects them all, so I can put a different tool in the reviewer pane than the one that wrote the code. A second model has different blind spots, and that’s the point of a second opinion. The template:

Scope: <path(s)> only. <Do not touch: path(s)>.
Context: Read <files / glossary / plan>. Use the terms <A, B, C>; never <overloaded term>.
Task (<verb: plan | implement | review>): <one sentence, one job>.
Constraints: <additive only | under N lines | no new dependencies | tests first>.
Output: <numbered list | file:line | severity first | markdown model | test names first>.
If anything above is ambiguous or blocks you, stop and ask.

That last line matters most. It turns “the agent is stuck” from a silent failure into a blocked state, which Herdr turns into a dot I can see from across the room.

7. When an agent goes blocked or done

The loop on my side is short and boring, on purpose:

  1. Blocked: read the question in place (it’s already on screen), answer it, and if the answer is a decision, write it into the notes repo’s decisions file so the next session doesn’t ask again.
  2. Done: read the output against the evidence pane next to it. If it’s a plan, approve it or correct it. If it’s findings, open a scoped implementation task in the other tab for anything worth fixing.
  3. Finished job: end that agent’s conversation. New job, new agent, clean context.

The point of all of it

The spaces, tabs, panes, and ultrawide exist to keep context small, boundaries visible, and stalls cheap. The prompting exists to keep meaning precise. Neither works without the other. A gorgeous Herdr layout full of agents with vague instructions is just a very wide way to watch things go wrong in parallel, and the most precise prompt in the world doesn’t help if the agent has been blocked on it since lunch and you never noticed.

Draw the boundaries on screen. Draw them again in words. Then let the agents work.

References

Herdr

Previously on Composite Code

Elsewhere

Read the whole story
alvinashcraft
6 minutes ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories