Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162531 stories
·
33 followers

python-1.21.0

1 Share

[1.21.0] - 2026-10-08

Added

  • agent-framework, agent-framework-core, agent-framework-purview: Add tool-result standing guidance, expose rewritten variable-expansion arguments, and support full-buffer Purview policy evaluation for streamed responses (#8784, #8506, #8702)
  • agent-framework-openai: Add the maximum reasoning-effort option (#8857)
  • agent-framework-oracle: Add the alpha Oracle native vector-store connector (#8676)
  • samples: Demonstrate group-chat response filtering and expand the get-started sample progression (#9005, #9133)

Changed

  • agent-framework-core, agent-framework-foundry, agent-framework-github-copilot: Isolate provider-owned service session state when agents are invoked as tools (#8853)
  • agent-framework-anthropic: Update the Anthropic SDK to 1.11 and adapt Anthropic and Bedrock client behavior (#9119)
  • agent-framework-github-copilot: Update the GitHub Copilot SDK to 1.0.16 (#9135)
  • agent-framework-devui: Update source-map-js, Radix UI Scroll Area, tailwind-merge, Tailwind CSS and its Vite plugin, and patch brace-expansion 1.x (#9104, #9154, #9155, #9156, #8888)
  • agent-framework-azure-contentunderstanding, agent-framework-azure-cosmos-memory, samples: Normalize Foundry project endpoint configuration and examples (#9070)
  • agent-framework-foundry: Clarify Foundry agent-tool ownership and lifecycle behavior (#8852)
  • agent-framework-foundry-hosting: Warn against hosting a WorkflowAgent through the agent= shortcut (#9146)
  • agent-framework-hyperlight, agent-framework-monty: Warn when CodeAct tools cannot support FIDES security enforcement (#9175)
  • agent-framework-hosting: Persist the session in the hosting quickstart (#9030)
  • docs: Add the AI-assisted contribution policy and issue-first contribution workflow (#9134)
  • samples: Update multidict and source-map-js lockfiles and make file-access samples type-check on Windows (#9105, #9145, #9161)
  • tests: Update Python GitHub Actions dependencies (#9038, #9039, #9158)
  • tests: Isolate the mixed-middleware warning assertion (#9001)
  • tests, samples: Refresh Python development dependencies and type-checker compatibility (#9190)

Fixed

  • agent-framework-core: Run local sibling tools in declaration-only mixed batches, serialize shared file edits, correct data URI validation and media detection, warn when streamed result gates lack buffering, and handle flat function mappings in progressive tool-name checks (#9050, #9032, #8918, #9029, #9117)
  • agent-framework-ag-ui, agent-framework-core: Close delegated provider streams when public streams are released while preserving primary failures during cleanup (#9100)
  • agent-framework-declarative: Preserve falsy values returned by Power Fx Search() (#8974)
  • agent-framework-devui: Page checkpoint item listings and scope function approvals to their conversations (#9126, #9180)
  • agent-framework-gemini, agent-framework-ollama: Close SDK streams when consumers stop early (#9144)
  • agent-framework-hosting: Avoid retaining per-session locks after runs complete (#9172)
  • agent-framework-hosting-responses: Parse input messages that omit the optional type field (#9124)
  • agent-framework-openai: Preserve streamed log probabilities across empty chunks (#8997)
  • agent-framework-orchestrations: Reject unknown group-chat participants and sanitize generated handoff tool names (#9026, #9035)
  • agent-framework-foundry-hosting: Preserve Foundry Toolbox container-file citations and refresh stale Toolbox call IDs (#9055, #9059)
  • agent-framework-typesafe: Honor proxy environment variables in connector-owned clients (#8984)
  • agent-framework-bedrock: Preserve extended-thinking reasoning content across tool calls (#8936)
  • agent-framework-tools: Preserve table-formatted output in persistent PowerShell sessions (#9043)

Full Changelog: python-1.20.0...python-1.21.0

Read the whole story
alvinashcraft
1 hour ago
reply
Pennsylvania, USA
Share this story
Delete

Introducing XP

1 Share

Our ambition at XBOX is to be the number one entertainment company in the world. This means growing our games into enduring franchises that entertain, inspire and connect people across every generation and across the globe.

That opportunity is why we’re creating XP.

As Asha shared in her email this morning, XP is a new XBOX division, led by newly appointed President, Kayleen Walters, spanning film and television, consumer products, strategic partnerships, and live experiences. It brings together capabilities that have existed across various teams at XBOX for years and aligns them behind a shared mission: helping our franchises reach more people, create deeper connections with players and new fans, and unlock new opportunities for these fandoms to grow around the world.

Great games remain at the center of these franchises and at the heart of XBOX. XP will partner with the best creative minds across our studios, along with trusted partners, to build on that foundation and create more ways for people to experience the worlds they love.

The formation of the XP division formalizes four investment areas:

  • XBOX PICTURES: Bringing our worlds and characters to new audiences through film, television, and other forms of linear storytelling.
  • XBOX PRODUCTS: Creating products that are authentic to the fandom that allow fans to celebrate and express the franchises they love.
  • XBOX PLACES: Bringing communities together through live events, immersive experiences, attractions, broadcasts, and destinations.
  • XBOX PARTNERSHIPS: Working with creators, brands, and communities to develop new experiences, sponsorships, and brand partnerships.

Throughout her career, Kayleen has helped steward some of the world’s most iconic entertainment properties. Across 13 years at Lucasfilm, she helped grow and relaunch one of the most influential franchises in history. At Mojang, her record speaks for itself. Under her tenure, Minecraft grew from an already-record setting 350 million units sold to over 425 million, its merchandising business became one of the top 20 globally. A Minecraft Movie had the biggest opening weekend ever for a video game adaptation and finished #1 at the domestic box office, and Mojang secured the first of many theme park placements to come.

Kayleen understands that the strongest franchises grow by remaining true to what and who made them special while creating new ways for fans, creators, partners, and communities to participate. I cannot imagine a better leader for this division and we are just getting started imagining what it will build.

Matthew Ball

The post Introducing XP appeared first on XBOX Wire.

Read the whole story
alvinashcraft
1 hour ago
reply
Pennsylvania, USA
Share this story
Delete

New Leadership Appointments and XBOX Division

1 Share

This message was just sent to Team XBOX employees globally.

Team,

At Town Hall on Tuesday, I talked about the four Cs and the choices we are making to put them into practice. You’ve seen a lot of our work on Core and Content. Today I want to share two leadership appointments supporting Creator and Connection. Kayleen Walters will transition from her role at Mojang to President of XP, a new XBOX division focused on how people experience our franchises beyond the games themselves, and Maria Angelidou-Smith will join XBOX as CEO of Mojang and Minecraft. Both will report to me.

Kayleen has spent her career building entertainment franchises. Before XBOX, she spent 13 years at Lucasfilm helping guide Star Wars across films, products, and partnerships. At Minecraft, she helped grow consumer products into a multibillion-dollar business with more than 250 licensees and was Executive Producer of A Minecraft Movie, the biggest film in the U.S. last year and a top-five release worldwide. She is also a producer of the sequel.

Image of Kayleen Walters, President of XP
Kayleen Walters, President of XP

XP is a play on “Experience Points” and formalizes four investment areas for us: XBOX PICTURES for film and television adaptations; XBOX PRODUCTS for licensing and merchandise; XBOX PARTNERSHIPS for sponsors and brand partners; and XBOX PLACES for our marquee broadcast and in-person events, theme parks, and other location-based experiences. Our franchises can mean more to people than the time they spend playing them, and Kayleen knows how to build those experiences with the industry without losing what made people care about the franchise in the first place.

Kayleen’s move also gave us time to think carefully about what Mojang needs in its next leader.

Minecraft is one of the largest communities in entertainment. Hundreds of millions of people play, build, create, teach, watch, mod, run servers, and build businesses around Minecraft. The community doesn’t just consume the product. It helps shape what the product becomes. That was central to our search. Mojang already has an exceptional game development team that knows Minecraft and its players deeply. We wanted someone who has built products for communities at enormous scale, who is deeply technical, and who understands how creators, product, technology, and monetization come together.  

Image of Maria Angelidou-Smith, CEO of Mojang and Minecraft
Maria Angelidou-Smith, CEO of Mojang and Minecraft

Maria brings a rare combination of those experiences. She spent nearly a decade at Meta, where she led Facebook Groups used by more than 1.8 billion people each month, as well as Events, Creators, Search, Profile, and Monetization. She later led product and technology at Personio and served as Chief Product Officer at Reddit, another of the world’s largest community platforms. She has built products used by billions of people and led product and engineering organizations through significant growth and change.

Maria brings deep experience to Mojang, along with a respect for what makes Minecraft special. Her first priority will be learning from the teams and communities who know Minecraft best. She starts Monday, October 19.

Please join me in congratulating Kayleen and welcoming Maria to XBOX!

Asha

The post New Leadership Appointments and XBOX Division appeared first on XBOX Wire.

Read the whole story
alvinashcraft
1 hour ago
reply
Pennsylvania, USA
Share this story
Delete

Arize's next top decision model

1 Share

Eight decision models. Three challenges. One bracket. Only one can be Arize’s next top decision model.

On 15 September, TypeSafe released Jev, and within a few weeks about 10 companies had shipped something that works the same way. OpenAI, Liquid, Cloudflare, AWS and PostHog all launched one, and an open clone called Kev showed up on Hugging Face in between. If you build agents or run evals, you now have a casting problem.

So we held auditions.

A decision model takes a question and a fixed set of options and returns a probability for each option. It doesn’t write text, so there are no output tokens to pay for or wait on. That makes it cheap and fast at two jobs AI engineers care about: making a single decision inside an agent (which tool, which skill, stop or carry on), and acting as the judge that scores your traces in Arize AX.

We’ve written about Jev as a judge before. Arize’s Head of Developer Relations Laurie Voss benchmarked it against LLM judges, and Elizabeth Hutton from the Arize Phoenix team looked at what its probabilities reveal. Both posts compared Jev with LLMs, but we hadn’t lined the new decision models up against each other on the same data. A spreadsheet of eight models is no fun to read, though, so we turned it into a knockout.

The best decision models are now level on accuracy. They differ on where you can run them, and on whether they understand the question the way you asked it.

Meet the contestants

Eight models made it through casting. Each one arrived with a different story, as all good contestants do.

Model Who Their story Size and licence Where we ran it List price
Jev 1.13 TypeSafe The one who started it all. Hosted API, 32k context Closed TypeSafe API $0.042 per million input tokens, output free
Liquid d1 Liquid AI First to knock Jev off the top of Hugging Face’s Decision Index, 58.9 to 57.9 Closed Liquid API $0.04 per million
OpenAI Decisions (gpt-6-luna) OpenAI The first frontier lab through the door. Still in public beta Closed OpenAI API $0.10 per million
Clef 27B Cloudflare Open weights, Jev-compatible API, can read images 27B, Apache 2.0 Workers AI $0.24 per million
Kev 9B Jared Palmer The underdog. An open Jev clone whose first training run, across three model sizes, cost about $95 of rented GPU time 9B, open Local Free to self-host
Jeeves 9B PostHog The overthinker. Reasons before it decides 9B, open Local Free to self-host
Strands Decider 2B AWS The smallest in the house, with its training recipe published 2B, open Local Free to self-host
Laya multilingual Convai Innovations Not even an LLM. A BERT-style encoder with an 8,192-token context 322M, Apache 2.0 Local Free to self-host

Prices are list prices per million input tokens. In the results, we turn them into cost per 1,000 decisions, using the tokens each model consumed.

Seven of the eight accept Jev’s /v1/systemone request format. Three weeks after launch, one API has become the default way to talk to a decision model, in the same way a lot of LLM vendors support the OpenAI API.

The odd one out is, fittingly, OpenAI. The other APIs take each case as named fields, such as the context and the response, but OpenAI’s Decisions API takes one block of plain text plus a list of questions. So we wrote each case out as labelled text sections, one per field, in the same order the other models saw them. We chose that layout, and a different one could score differently. Treat OpenAI’s numbers as a fair first look at a public beta, and a better prompt layout might lift them.

The rules of the house

The bracket: eight entrants in two conferences

Comparing a hosted API with a model running on a laptop isn’t a fair fight on speed, so the house is split into two conferences.

  • The cloud conference is hosted APIs, called over the network.
  • The local conference is open models running on a MacBook Pro (Apple M4 Max, 36 GB of memory).

Each conference crowns a champion, and the two meet in the grand final.

There are three challenges, one per round:

  • The quarterfinals are the hallucination challenge: 2,675 rows from RAGTruth, built with our own benchmark code.
  • The conference finals are the evaluator challenge: all 13 Phoenix evaluator suites, 655 labelled cases covering hallucination, toxicity, tool use and more.
  • The grand final is the agent challenge: 52 routing decisions where the model picks which tool or skill an agent should use.

The judging panel scores three things: accuracy, latency and cost per 1,000 decisions. Win two of the three and you go through. Accuracy only counts as a win when the 95% confidence intervals don’t overlap; otherwise that category is a draw. Ties go to accuracy, or to calibration in the conference finals. If that can’t separate them either, the match is a draw. The grand final drops latency, because a laptop and an API still aren’t comparable.

To keep it fair, every model gets its own yes/no cutoff, tuned on one half of the data and scored on the other half. It’s the same procedure and random seed we’ve used in the past.

Quarterfinals: the hallucination challenge

The bracket after the quarterfinals

The first challenge is one of the more popular evals, hallucination detection. Given a context and a response, is every claim in the response supported by the context? RAGTruth mixes question answering, summaries and data-to-text, and about a third of its responses contain a hallucination.

Jev vs OpenAI. The first frontier lab in the house drew the house favourite in round one. OpenAI scored 81.5% balanced accuracy against Jev’s 85.5%. That looks like a clear gap, but the confidence intervals touch by a tenth of a point, so under our rules accuracy is a draw. The other two categories weren’t close. Jev answered in 0.08 seconds per decision against 0.16, and cost $0.045 per 1,000 decisions against $0.090. Jev goes through 2–0, and OpenAI packs its bags early. Don’t forget about it, though. It comes back in the twist.

Liquid d1 vs Clef 27B. This is the match the panel will be talking about for weeks. On accuracy it was a dead heat: Clef 27B scored 85.2%, Liquid 84.8%, well inside each other’s confidence intervals. So the decision came down to the other two categories, and Liquid won both. It was faster (0.20 seconds against 0.50) and much cheaper ($0.034 per 1,000 against $0.223). Clef 27B, I’m sorry, but you’re going home.

Kev vs Laya. Laya was the fastest model in the competition at 0.04 seconds a decision, and it’s under a gigabyte. But it scored 59.8% to Kev’s 76.8%. On a laptop the cost is the same for both, so accuracy decides. Kev goes through.

Jeeves vs Strands. Jeeves reasons before it answers, and on hallucination that thinking paid off: 80.3% against 69.0% for Strands. Strands answered about 50 times faster at the median, but speed alone can’t win a match where cost is tied. Jeeves goes through, slowly.

How do the survivors compare with an LLM judge? In our earlier benchmarks, Opus 5 scored 86.6% balanced accuracy at $14.30 per 1,000 judgments. Jev, Liquid and Clef 27B all land within the noise of that at a fraction of the price: Jev costs 0.3% of what Opus 5 does, Liquid 0.2% and Clef 27B 1.6%. OpenAI sits a few points further back, at 0.6%.

Conference finals: the evaluator challenge

The bracket after the conference finals

This round is the audition for the job most of you are hiring for: could this model be the judge in your eval pipeline? The 13 Phoenix suites test the evaluators people run in production, from faithfulness and conciseness to whether an agent handled a tool response correctly.

It’s also where the panel brings out its favourite tiebreaker: calibration. A calibrated model is right 99% of the time when it says it’s 99% sure, and 70% of the time when it says it’s 70% sure. For a judge it counts as much as raw accuracy, because it tells you which verdicts you can trust and which ones should go to a human or a bigger model.

We score it as calibration error: the average gap between how sure a model says it is and how often it’s right, so lower is better. We use the same calibration code as our earlier Jev benchmark, with every model’s confidence on the same scale.

Cloud final: Jev vs Liquid. Jev scored 91.1% and Liquid 87.6%, and those intervals overlap, so accuracy is a draw. Jev won on speed and Liquid on cost. That’s one each, so the tie goes to calibration. Jev’s calibration error was 0.023 and Liquid’s 0.051, so on average Jev’s confidence sat within 2 points of how often it was right, and Liquid’s within 5. Both are excellent, and with 655 cases their confidence intervals overlap. The panel can’t split them. The cloud final is a draw.

If you made us pick, Jev edges it by a whisker. Its confidence was a slightly better guide to its own mistakes: take one case it got right and one it got wrong, and 85% of the time it was more confident about the right one, against 81% for Liquid. On the 14% of cases where Jev said it was at least 99% sure, it was right every time. Liquid was 99% sure on a third of the cases and right 99.1% of the time, which is also very good. Jev goes through to the grand final, by a nose.

That tells you something about the hosted APIs: at the top, they’re close to interchangeable.

Local final: Kev vs Jeeves. Kev scored 88.7%, Jeeves 87.1%. Another draw on accuracy, and a draw on cost. But Kev answered in 0.19 seconds and Jeeves in 4.3, more than 20 times slower. On this challenge, all that thinking didn’t buy Jeeves any accuracy. Kev didn’t need the tiebreak, but for the record the two were level on calibration, 0.055 against 0.043 for Jeeves, well inside the noise. Kev’s confidence was the better guide to its own mistakes, 84% against 79% on the same test as the cloud final. The underdog is your local champion.

Grand final: the agent challenge

The final bracket: Jev and Kev draw

The last challenge leaves the judging booth and steps inside an agent. Each case is a user message and a list of tools or skills, and the model has to pick the right one, including “none of them”. Twenty cases come from the travel assistant in our tool-calling evaluation post, and 32 from the labelled routing tests for Phoenix’s in-app assistant.

Jev got 49 of 52 right. Kev got 48. With only 52 cases, one question is worth about two percentage points, so a single answer is well inside the noise. Accuracy is a draw.

Cost can’t separate them fairly either. Jev costs $0.022 per 1,000 routing decisions. Kev costs nothing extra if you already own the hardware, and rather more if you have to rent a GPU to run it.

I have two photos in my hand. And I’m handing out both of them.

The grand final is a draw. If you want an API you call and forget, Jev is Arize’s next top decision model. If you want to self-host, Kev is. The scores can’t split them, so your preference for where to run picks the winner.

The twist: ask the question the wrong way and most models flip

Surprisingly, the biggest gap in the competition isn’t in the bracket.

Before the runs, we tried a few ways of wording the hallucination question. Our original RAG test prompt asks whether the response contains an unsupported claim. Our final question asks whether every claim is supported. They mean the same thing with the polarity reversed, so a model that understands the question should score about the same on both.

So we ran every model a second time with the “unsupported claim” wording, on the same 1,338 test rows. Here’s ROC AUC, where 0.5 is a coin flip and anything under 0.5 means the model is answering the opposite question:

Model “Is every claim supported?” “Does it contain an unsupported claim?”
Jev 1.13 0.93 0.93
Clef 27B 0.92 0.91
OpenAI Decisions 0.90 0.89
Kev 9B 0.85 0.54
Laya multilingual 0.64 0.36
Strands Decider 2B 0.75 0.26
Jeeves 9B 0.85 0.25
Liquid d1 0.92 0.17

Only Jev, Clef 27B and OpenAI held their pose. OpenAI went home in round one, but it’s one of only three models that understood the question both ways. Kev dropped to a coin flip. Liquid, Jeeves and Strands flipped, confidently answering “is this response grounded?” when we asked “does it contain an unsupported claim?” Liquid went from one of the best scores in the competition to 0.17.

Laya showed the same weakness on everyday phrasing. On the Phoenix suites worded as “Is the text free of toxic content?”, it scored 0.11. Flip the wording to “Does the text contain toxic content?” and it jumps to 0.87. Its model card warns that its answers can follow the option labels rather than the question.

A decision model’s score belongs to the model and the question. Swap a new model into your judge without changing a word, and you may have inverted your evals without noticing. The fix is cheap: run the new model on both phrasings of your question before you trust it.

The one who should have made the final: Clef 27B

Every season has a contestant who goes home too early, and the internet doesn’t let it go. This season, it’s Clef 27B.

Model Hallucination (RAGTruth) Evaluators (Phoenix suites) Routing Calibration error Latency (p50, RAGTruth) Cost per 1,000 (RAGTruth)
Jev 1.13 85.5% 91.1% 49/52 0.023 0.08 s $0.045
Clef 27B 85.2% 91.5% 50/52 0.062 0.50 s $0.223
Liquid d1 84.8% 87.6% 48/52 0.051 0.20 s $0.034
OpenAI Decisions 81.5% 90.0% 48/52 0.042 0.16 s $0.090
Jeeves 9B 80.3% 87.1% 48/52 0.043 16.8 s $0 (local)
Kev 9B 76.8% 88.7% 48/52 0.055 1.03 s $0 (local)
Strands Decider 2B 69.0% 76.7% 45/52 0.065 0.34 s $0 (local)
Laya multilingual 59.8% 57.3% 40/52 0.330 0.04 s $0 (local)

Balanced accuracy on the held-out half at each model’s tuned cutoff. Routing is all 52 cases. Calibration error is on all 655 Phoenix cases, with every model’s confidence on the same scale (lower is better). Local latencies are from one laptop and aren’t comparable with the cloud numbers.

Clef 27B matched Jev on accuracy in every round, and along with Jev and OpenAI it was one of only three models to survive the wording twist. It went out in round one because Liquid was faster and cheaper on the day, and against Jev it would have lost on those same two categories.

That’s the trouble with knockouts, and why we published the full results. If you want open weights you can also call as a hosted API, Clef 27B is the one to look at.

The final walk

Three weeks after Jev, the top decision models are a commodity on accuracy. Jev, Clef 27B and Liquid sit within a point of each other on hallucination, and within the noise of an LLM judge that costs 60 to 400 times more. Even the cloud final couldn’t separate Jev and Liquid. Kev, an open clone whose first training run cost about $95 in rented GPUs, ties Jev in the agent challenge.

Pick on where you want to run it, an API you call or weights you host. Then check the model reads your question the way you meant it, because wording can flip a judge from right to confidently wrong.

Test the question as hard as you test the model.

Your turn on the runway

You don’t need to rebuild our harness to run your own season. Arize AX has Jev as a judge built in: add your TypeSafe API key, write your question, and Jev scores your traces with a label and a confidence. If your contestant is one we didn’t bring, like Kev on your own GPU, put it behind an HTTPS endpoint and connect it as a remote evaluator. Your model does the scoring, and AX handles calling it, retries and writing the results back onto your traces. Either way, put both phrasings of your question to it before you trust the scores.

Read the whole story
alvinashcraft
1 hour ago
reply
Pennsylvania, USA
Share this story
Delete

Real-world lessons in agentic authority and overreach

1 Share

Real-world lessons in agentic authority and overreach

Read the whole story
alvinashcraft
1 hour ago
reply
Pennsylvania, USA
Share this story
Delete

Claude Haiku 5.5 Pricing: The 100K-Token Rule for Agents

1 Share

Claude Haiku 5.5 costs $0.10 per million input tokens up to 100K tokens of prompt, and $0.50 per million above it, a fivefold cliff that decides your agent bill. Anthropic released the model on 7 October 2026 on its own platform, AWS, Google Cloud and Microsoft Azure the same day. Output is $0.50 per million tokens at or below the threshold and $2.50 above it. The context window is 1M tokens with up to 128K tokens of output.

Short answer: for high-volume agent loops, Haiku 5.5 is the cheapest Anthropic model ever listed, and a loop that keeps every prompt under 100K tokens runs about $1,050 per million calls at 8,000 input tokens and 500 output tokens each. Add prompt caching and that falls to roughly $520. The risk is the 100K line. One careless context-stuffing change moves a call from the cheap tier to the expensive one. Here is the price sheet, the worked budget and the rules that keep you below the line.

The two-tier price sheet

The numbers below come from Anthropic's pricing page as reported by VentureBeat and MarkTechPost on launch day. Anthropic describes Haiku 5.5 as its "best lightweight model yet" and says it beats Haiku 4.5 in coding, computer use and knowledge work. It also says the model is on average 75% cheaper than Haiku 4.5. VentureBeat frames the headline as a 90% API price cut against Haiku 4.5's $1.00 input and $5.00 output, which is what you get when you compare the lower tier alone.

Item Prompt up to 100K tokens Prompt above 100K tokens

| Input, per million tokens | $0.10 | $0.50 |

| Output, per million tokens | $0.50 | $2.50 |

| Cache read, per million tokens | $0.01 | not listed in my sources |

| Cache write (5 minute), per million tokens | $0.125 | not listed in my sources |

Read the two columns as a step function, not a slope. A 99,000-token prompt with a 1,000-token answer costs $0.0104. A 101,000-token prompt with the same answer costs $0.053, which is 5.1 times more for 2% more input. That calculation assumes the higher rate applies to the whole request, which is how the tier is listed. Confirm on Anthropic's pricing page whether your region and provider bill it that way before you build a forecast on it.

The cache read price is worth a second look: $0.01 per million is one tenth of the base input price. If most of your prompt is a fixed system prompt, tool definitions and few-shot examples, you pay a tenth for those tokens on every hit. The cache write costs $0.125 per million, a 25% premium over base input, charged when the prefix is first stored.

A worked budget for one million calls

Assume a classification or routing agent loop. Each call sends 8,000 input tokens and returns 500 output tokens. Of the input, 6,000 tokens are a fixed prefix: instructions, tool schemas and examples. The other 2,000 vary per call. All numbers are mine, for illustration. Swap in your own.

Scenario (1M calls, 8,000 in, 500 out) Input cost Output cost Total

| Haiku 5.5, no caching | 8,000M tokens x $0.10 = $800 | 500M x $0.50 = $250 | $1,050 |

| Haiku 5.5, 6,000-token prefix cached | 6,000M x $0.01 + 2,000M x $0.10 = $260 | $250 | $510, plus about $7.50 of cache writes |

| Haiku 4.5 ($1.00 in, $5.00 out) | $8,000 | $2,500 | $10,500 |

| Gemini 3.5 Flash ($1.50 in, $9.00 out) | $12,000 | $4,500 | $16,500 |

The cache write line assumes one prefix write per 100 calls, which is 10,000 writes of 6,000 tokens at $0.125 per million. That is a hot-cache assumption. If your traffic is bursty and the five-minute cache expires between bursts, you pay the write more often, and the saving shrinks.

The ratio is the point. Under identical traffic Haiku 5.5 comes out near one tenth of Haiku 4.5 and under one fifteenth of Gemini 3.5 Flash at its listed $1.50 and $9.00. Gemini 3.5 Flash launched on 19 May 2026 with a cached price of $0.15 and batch pricing of $0.75 and $4.50, so its gap narrows with batch jobs but does not close. A newer Gemini 3.8 Flash is reported at an introductory $0.75 and $3.75 through 31 December 2026, then $1.50 and $7.50. Sources disagree on those figures, so check Google's page before relying on them.

OpenAI's GPT-6 Luna is the direct competitor. VentureBeat describes Haiku 5.5 as priced in line with it. I do not have a verified Luna price sheet from this week, so I am not putting a Luna row in the table. Price your own workload on the AI Model Cost Calculator with both models side by side.

Staying under the 100K line

The 1M context window is a capability, not a budget. Using it flips every call into the higher tier. Three habits keep an agent loop in the cheap lane.

First, cap the prompt in code, not in hope. Count tokens before you send. If a request crosses 90,000 tokens, summarise or drop the oldest turns instead of sending it. A ten percent safety margin costs you little and removes the cliff from your daily risk.

Second, treat tool output as the usual culprit. Agents that read files, web pages or search results grow their context fastest through tool results. Truncate results at the tool boundary and return identifiers the agent can fetch on demand, not whole documents.

Third, route the rare big job elsewhere on purpose. If a task genuinely needs 400K tokens of context, it is a different workload. Price it at the higher tier deliberately or give it to a model built for long inputs. Do not let it arrive by accident through a loop that never trims history.

// guard before every call (pseudo-code)
const tokens = countTokens(messages)
if (tokens > 90_000) messages = compact(messages)  // summarise old turns
send(messages)

The new Agent Run Cost Simulator models a multi-step run, not a single call, which is where the tier line bites: step ten of an agent run carries the history of steps one to nine. Run your expected step count through it with the 100K threshold in mind.

What the cheap tier does not fix

Four costs survive the price cut, and they are the ones that surprise teams after the first invoice.

Retries multiply everything. A loop that retries a failed step three times triples the spend on that step, and an agent that wanders can retry silently. Put a hard cap on steps and attempts per run, and log the count per run so you can see the tail. A run that costs 40 times the median is a bug, and it hides inside an average.

Output tokens cost five times input at the lower tier. At $0.50 against $0.10 per million, a chatty model is expensive in a way a terse one is not. In the budget above, output was 24% of the uncached bill but 49% of the cached one. Once you cache the prefix, output is where the money goes, so ask for structured, short answers and set a maximum output length on every call.

Cache misses are silent. The cache read price is one tenth of base input, but only a hit earns it. Change one character early in the prefix, reorder your tool definitions, or insert a timestamp near the top, and the next call pays full price and pays a write on top. Keep the stable part first and the volatile part last, and check the cache hit numbers your provider returns on each response.

Rate limits are a separate ceiling. A million calls a day is about 12 calls a second sustained. Your account limits, not the price, may decide how fast you can run that. Check the limits for your tier before you promise a batch finishes by morning.

None of this argues against the model. It argues for measuring one thing before you scale: cost per successful task, not cost per call. Divide your total spend by the number of tasks that finished correctly. That figure is the one that stays honest when prices, models and prompts change again next month.

Where Haiku 5.5 fits in an agent stack

A model this cheap changes the architecture question from "can I afford a call here" to "which calls do I still send to a bigger model". Haiku 5.5 is a fit for routing and triage, extraction from structured text, log and ticket classification, first-pass code review comments, and the many small steps in a computer-use or browser agent. It is the wrong default for the single hard step that decides whether a run succeeds. For that step, pay for a stronger model and keep Haiku on everything around it.

The honest limit: Anthropic's claim that Haiku 5.5 beats Haiku 4.5 is a vendor claim, and a lower price tells you nothing about your accuracy. Run 200 of your real inputs through both models before you switch, and compare the failures, not the average score. A cheaper model that needs a retry on 15% of calls is not 90% cheaper.

Cost control is the discipline that makes cheap models safe. The AI Agent Ops Bundle covers specs, observability and cost control for exactly this kind of loop, and the Agent Prompt Vault gives you 50 production prompts with stable prefixes you can cache. Fleets of agents also have a coordination bill, which the post The Multi-Agent Tax covers. If your volume is on the OpenAI side, the sibling post on Codex Cloud pricing shows how a rate limit differs from a rate.

Quick answers

How much does Claude Haiku 5.5 cost?

For prompts up to 100K tokens it is $0.10 per million input tokens and $0.50 per million output tokens. Above 100K tokens it is $0.50 input and $2.50 output. Cache reads are $0.01 per million and 5-minute cache writes are $0.125 per million.

What is the Haiku 5.5 context window?

One million tokens, with up to 128K tokens of output. Using more than 100K tokens of prompt moves the call to the higher price tier.

Is Haiku 5.5 really 90% cheaper than Haiku 4.5?

At the lower tier, yes: $0.10 and $0.50 against $1.00 and $5.00. Anthropic's own average figure is 75% cheaper, which is the number to use if some of your calls cross the 100K threshold.

Which models compete with Haiku 5.5 on price?

VentureBeat says OpenAI GPT-6 Luna is priced in line with it. Gemini 3.5 Flash lists at $1.50 input and $9.00 output, so it costs about 15 times more at the same token mix.

Every product mentioned is available at wowhow.cloud — pay once, ship forever.

Originally published at wowhow.cloud

Read the whole story
alvinashcraft
1 hour ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories