Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162528 stories
·
33 followers

Arize's next top decision model

1 Share

Eight decision models. Three challenges. One bracket. Only one can be Arize’s next top decision model.

On 15 September, TypeSafe released Jev, and within a few weeks about 10 companies had shipped something that works the same way. OpenAI, Liquid, Cloudflare, AWS and PostHog all launched one, and an open clone called Kev showed up on Hugging Face in between. If you build agents or run evals, you now have a casting problem.

So we held auditions.

A decision model takes a question and a fixed set of options and returns a probability for each option. It doesn’t write text, so there are no output tokens to pay for or wait on. That makes it cheap and fast at two jobs AI engineers care about: making a single decision inside an agent (which tool, which skill, stop or carry on), and acting as the judge that scores your traces in Arize AX.

We’ve written about Jev as a judge before. Arize’s Head of Developer Relations Laurie Voss benchmarked it against LLM judges, and Elizabeth Hutton from the Arize Phoenix team looked at what its probabilities reveal. Both posts compared Jev with LLMs, but we hadn’t lined the new decision models up against each other on the same data. A spreadsheet of eight models is no fun to read, though, so we turned it into a knockout.

The best decision models are now level on accuracy. They differ on where you can run them, and on whether they understand the question the way you asked it.

Meet the contestants

Eight models made it through casting. Each one arrived with a different story, as all good contestants do.

Model Who Their story Size and licence Where we ran it List price
Jev 1.13 TypeSafe The one who started it all. Hosted API, 32k context Closed TypeSafe API $0.042 per million input tokens, output free
Liquid d1 Liquid AI First to knock Jev off the top of Hugging Face’s Decision Index, 58.9 to 57.9 Closed Liquid API $0.04 per million
OpenAI Decisions (gpt-6-luna) OpenAI The first frontier lab through the door. Still in public beta Closed OpenAI API $0.10 per million
Clef 27B Cloudflare Open weights, Jev-compatible API, can read images 27B, Apache 2.0 Workers AI $0.24 per million
Kev 9B Jared Palmer The underdog. An open Jev clone whose first training run, across three model sizes, cost about $95 of rented GPU time 9B, open Local Free to self-host
Jeeves 9B PostHog The overthinker. Reasons before it decides 9B, open Local Free to self-host
Strands Decider 2B AWS The smallest in the house, with its training recipe published 2B, open Local Free to self-host
Laya multilingual Convai Innovations Not even an LLM. A BERT-style encoder with an 8,192-token context 322M, Apache 2.0 Local Free to self-host

Prices are list prices per million input tokens. In the results, we turn them into cost per 1,000 decisions, using the tokens each model consumed.

Seven of the eight accept Jev’s /v1/systemone request format. Three weeks after launch, one API has become the default way to talk to a decision model, in the same way a lot of LLM vendors support the OpenAI API.

The odd one out is, fittingly, OpenAI. The other APIs take each case as named fields, such as the context and the response, but OpenAI’s Decisions API takes one block of plain text plus a list of questions. So we wrote each case out as labelled text sections, one per field, in the same order the other models saw them. We chose that layout, and a different one could score differently. Treat OpenAI’s numbers as a fair first look at a public beta, and a better prompt layout might lift them.

The rules of the house

The bracket: eight entrants in two conferences

Comparing a hosted API with a model running on a laptop isn’t a fair fight on speed, so the house is split into two conferences.

  • The cloud conference is hosted APIs, called over the network.
  • The local conference is open models running on a MacBook Pro (Apple M4 Max, 36 GB of memory).

Each conference crowns a champion, and the two meet in the grand final.

There are three challenges, one per round:

  • The quarterfinals are the hallucination challenge: 2,675 rows from RAGTruth, built with our own benchmark code.
  • The conference finals are the evaluator challenge: all 13 Phoenix evaluator suites, 655 labelled cases covering hallucination, toxicity, tool use and more.
  • The grand final is the agent challenge: 52 routing decisions where the model picks which tool or skill an agent should use.

The judging panel scores three things: accuracy, latency and cost per 1,000 decisions. Win two of the three and you go through. Accuracy only counts as a win when the 95% confidence intervals don’t overlap; otherwise that category is a draw. Ties go to accuracy, or to calibration in the conference finals. If that can’t separate them either, the match is a draw. The grand final drops latency, because a laptop and an API still aren’t comparable.

To keep it fair, every model gets its own yes/no cutoff, tuned on one half of the data and scored on the other half. It’s the same procedure and random seed we’ve used in the past.

Quarterfinals: the hallucination challenge

The bracket after the quarterfinals

The first challenge is one of the more popular evals, hallucination detection. Given a context and a response, is every claim in the response supported by the context? RAGTruth mixes question answering, summaries and data-to-text, and about a third of its responses contain a hallucination.

Jev vs OpenAI. The first frontier lab in the house drew the house favourite in round one. OpenAI scored 81.5% balanced accuracy against Jev’s 85.5%. That looks like a clear gap, but the confidence intervals touch by a tenth of a point, so under our rules accuracy is a draw. The other two categories weren’t close. Jev answered in 0.08 seconds per decision against 0.16, and cost $0.045 per 1,000 decisions against $0.090. Jev goes through 2–0, and OpenAI packs its bags early. Don’t forget about it, though. It comes back in the twist.

Liquid d1 vs Clef 27B. This is the match the panel will be talking about for weeks. On accuracy it was a dead heat: Clef 27B scored 85.2%, Liquid 84.8%, well inside each other’s confidence intervals. So the decision came down to the other two categories, and Liquid won both. It was faster (0.20 seconds against 0.50) and much cheaper ($0.034 per 1,000 against $0.223). Clef 27B, I’m sorry, but you’re going home.

Kev vs Laya. Laya was the fastest model in the competition at 0.04 seconds a decision, and it’s under a gigabyte. But it scored 59.8% to Kev’s 76.8%. On a laptop the cost is the same for both, so accuracy decides. Kev goes through.

Jeeves vs Strands. Jeeves reasons before it answers, and on hallucination that thinking paid off: 80.3% against 69.0% for Strands. Strands answered about 50 times faster at the median, but speed alone can’t win a match where cost is tied. Jeeves goes through, slowly.

How do the survivors compare with an LLM judge? In our earlier benchmarks, Opus 5 scored 86.6% balanced accuracy at $14.30 per 1,000 judgments. Jev, Liquid and Clef 27B all land within the noise of that at a fraction of the price: Jev costs 0.3% of what Opus 5 does, Liquid 0.2% and Clef 27B 1.6%. OpenAI sits a few points further back, at 0.6%.

Conference finals: the evaluator challenge

The bracket after the conference finals

This round is the audition for the job most of you are hiring for: could this model be the judge in your eval pipeline? The 13 Phoenix suites test the evaluators people run in production, from faithfulness and conciseness to whether an agent handled a tool response correctly.

It’s also where the panel brings out its favourite tiebreaker: calibration. A calibrated model is right 99% of the time when it says it’s 99% sure, and 70% of the time when it says it’s 70% sure. For a judge it counts as much as raw accuracy, because it tells you which verdicts you can trust and which ones should go to a human or a bigger model.

We score it as calibration error: the average gap between how sure a model says it is and how often it’s right, so lower is better. We use the same calibration code as our earlier Jev benchmark, with every model’s confidence on the same scale.

Cloud final: Jev vs Liquid. Jev scored 91.1% and Liquid 87.6%, and those intervals overlap, so accuracy is a draw. Jev won on speed and Liquid on cost. That’s one each, so the tie goes to calibration. Jev’s calibration error was 0.023 and Liquid’s 0.051, so on average Jev’s confidence sat within 2 points of how often it was right, and Liquid’s within 5. Both are excellent, and with 655 cases their confidence intervals overlap. The panel can’t split them. The cloud final is a draw.

If you made us pick, Jev edges it by a whisker. Its confidence was a slightly better guide to its own mistakes: take one case it got right and one it got wrong, and 85% of the time it was more confident about the right one, against 81% for Liquid. On the 14% of cases where Jev said it was at least 99% sure, it was right every time. Liquid was 99% sure on a third of the cases and right 99.1% of the time, which is also very good. Jev goes through to the grand final, by a nose.

That tells you something about the hosted APIs: at the top, they’re close to interchangeable.

Local final: Kev vs Jeeves. Kev scored 88.7%, Jeeves 87.1%. Another draw on accuracy, and a draw on cost. But Kev answered in 0.19 seconds and Jeeves in 4.3, more than 20 times slower. On this challenge, all that thinking didn’t buy Jeeves any accuracy. Kev didn’t need the tiebreak, but for the record the two were level on calibration, 0.055 against 0.043 for Jeeves, well inside the noise. Kev’s confidence was the better guide to its own mistakes, 84% against 79% on the same test as the cloud final. The underdog is your local champion.

Grand final: the agent challenge

The final bracket: Jev and Kev draw

The last challenge leaves the judging booth and steps inside an agent. Each case is a user message and a list of tools or skills, and the model has to pick the right one, including “none of them”. Twenty cases come from the travel assistant in our tool-calling evaluation post, and 32 from the labelled routing tests for Phoenix’s in-app assistant.

Jev got 49 of 52 right. Kev got 48. With only 52 cases, one question is worth about two percentage points, so a single answer is well inside the noise. Accuracy is a draw.

Cost can’t separate them fairly either. Jev costs $0.022 per 1,000 routing decisions. Kev costs nothing extra if you already own the hardware, and rather more if you have to rent a GPU to run it.

I have two photos in my hand. And I’m handing out both of them.

The grand final is a draw. If you want an API you call and forget, Jev is Arize’s next top decision model. If you want to self-host, Kev is. The scores can’t split them, so your preference for where to run picks the winner.

The twist: ask the question the wrong way and most models flip

Surprisingly, the biggest gap in the competition isn’t in the bracket.

Before the runs, we tried a few ways of wording the hallucination question. Our original RAG test prompt asks whether the response contains an unsupported claim. Our final question asks whether every claim is supported. They mean the same thing with the polarity reversed, so a model that understands the question should score about the same on both.

So we ran every model a second time with the “unsupported claim” wording, on the same 1,338 test rows. Here’s ROC AUC, where 0.5 is a coin flip and anything under 0.5 means the model is answering the opposite question:

Model “Is every claim supported?” “Does it contain an unsupported claim?”
Jev 1.13 0.93 0.93
Clef 27B 0.92 0.91
OpenAI Decisions 0.90 0.89
Kev 9B 0.85 0.54
Laya multilingual 0.64 0.36
Strands Decider 2B 0.75 0.26
Jeeves 9B 0.85 0.25
Liquid d1 0.92 0.17

Only Jev, Clef 27B and OpenAI held their pose. OpenAI went home in round one, but it’s one of only three models that understood the question both ways. Kev dropped to a coin flip. Liquid, Jeeves and Strands flipped, confidently answering “is this response grounded?” when we asked “does it contain an unsupported claim?” Liquid went from one of the best scores in the competition to 0.17.

Laya showed the same weakness on everyday phrasing. On the Phoenix suites worded as “Is the text free of toxic content?”, it scored 0.11. Flip the wording to “Does the text contain toxic content?” and it jumps to 0.87. Its model card warns that its answers can follow the option labels rather than the question.

A decision model’s score belongs to the model and the question. Swap a new model into your judge without changing a word, and you may have inverted your evals without noticing. The fix is cheap: run the new model on both phrasings of your question before you trust it.

The one who should have made the final: Clef 27B

Every season has a contestant who goes home too early, and the internet doesn’t let it go. This season, it’s Clef 27B.

Model Hallucination (RAGTruth) Evaluators (Phoenix suites) Routing Calibration error Latency (p50, RAGTruth) Cost per 1,000 (RAGTruth)
Jev 1.13 85.5% 91.1% 49/52 0.023 0.08 s $0.045
Clef 27B 85.2% 91.5% 50/52 0.062 0.50 s $0.223
Liquid d1 84.8% 87.6% 48/52 0.051 0.20 s $0.034
OpenAI Decisions 81.5% 90.0% 48/52 0.042 0.16 s $0.090
Jeeves 9B 80.3% 87.1% 48/52 0.043 16.8 s $0 (local)
Kev 9B 76.8% 88.7% 48/52 0.055 1.03 s $0 (local)
Strands Decider 2B 69.0% 76.7% 45/52 0.065 0.34 s $0 (local)
Laya multilingual 59.8% 57.3% 40/52 0.330 0.04 s $0 (local)

Balanced accuracy on the held-out half at each model’s tuned cutoff. Routing is all 52 cases. Calibration error is on all 655 Phoenix cases, with every model’s confidence on the same scale (lower is better). Local latencies are from one laptop and aren’t comparable with the cloud numbers.

Clef 27B matched Jev on accuracy in every round, and along with Jev and OpenAI it was one of only three models to survive the wording twist. It went out in round one because Liquid was faster and cheaper on the day, and against Jev it would have lost on those same two categories.

That’s the trouble with knockouts, and why we published the full results. If you want open weights you can also call as a hosted API, Clef 27B is the one to look at.

The final walk

Three weeks after Jev, the top decision models are a commodity on accuracy. Jev, Clef 27B and Liquid sit within a point of each other on hallucination, and within the noise of an LLM judge that costs 60 to 400 times more. Even the cloud final couldn’t separate Jev and Liquid. Kev, an open clone whose first training run cost about $95 in rented GPUs, ties Jev in the agent challenge.

Pick on where you want to run it, an API you call or weights you host. Then check the model reads your question the way you meant it, because wording can flip a judge from right to confidently wrong.

Test the question as hard as you test the model.

Your turn on the runway

You don’t need to rebuild our harness to run your own season. Arize AX has Jev as a judge built in: add your TypeSafe API key, write your question, and Jev scores your traces with a label and a confidence. If your contestant is one we didn’t bring, like Kev on your own GPU, put it behind an HTTPS endpoint and connect it as a remote evaluator. Your model does the scoring, and AX handles calling it, retries and writing the results back onto your traces. Either way, put both phrasings of your question to it before you trust the scores.

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Real-world lessons in agentic authority and overreach

1 Share

Real-world lessons in agentic authority and overreach

Read the whole story
alvinashcraft
28 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Claude Haiku 5.5 Pricing: The 100K-Token Rule for Agents

1 Share

Claude Haiku 5.5 costs $0.10 per million input tokens up to 100K tokens of prompt, and $0.50 per million above it, a fivefold cliff that decides your agent bill. Anthropic released the model on 7 October 2026 on its own platform, AWS, Google Cloud and Microsoft Azure the same day. Output is $0.50 per million tokens at or below the threshold and $2.50 above it. The context window is 1M tokens with up to 128K tokens of output.

Short answer: for high-volume agent loops, Haiku 5.5 is the cheapest Anthropic model ever listed, and a loop that keeps every prompt under 100K tokens runs about $1,050 per million calls at 8,000 input tokens and 500 output tokens each. Add prompt caching and that falls to roughly $520. The risk is the 100K line. One careless context-stuffing change moves a call from the cheap tier to the expensive one. Here is the price sheet, the worked budget and the rules that keep you below the line.

The two-tier price sheet

The numbers below come from Anthropic's pricing page as reported by VentureBeat and MarkTechPost on launch day. Anthropic describes Haiku 5.5 as its "best lightweight model yet" and says it beats Haiku 4.5 in coding, computer use and knowledge work. It also says the model is on average 75% cheaper than Haiku 4.5. VentureBeat frames the headline as a 90% API price cut against Haiku 4.5's $1.00 input and $5.00 output, which is what you get when you compare the lower tier alone.

Item Prompt up to 100K tokens Prompt above 100K tokens

| Input, per million tokens | $0.10 | $0.50 |

| Output, per million tokens | $0.50 | $2.50 |

| Cache read, per million tokens | $0.01 | not listed in my sources |

| Cache write (5 minute), per million tokens | $0.125 | not listed in my sources |

Read the two columns as a step function, not a slope. A 99,000-token prompt with a 1,000-token answer costs $0.0104. A 101,000-token prompt with the same answer costs $0.053, which is 5.1 times more for 2% more input. That calculation assumes the higher rate applies to the whole request, which is how the tier is listed. Confirm on Anthropic's pricing page whether your region and provider bill it that way before you build a forecast on it.

The cache read price is worth a second look: $0.01 per million is one tenth of the base input price. If most of your prompt is a fixed system prompt, tool definitions and few-shot examples, you pay a tenth for those tokens on every hit. The cache write costs $0.125 per million, a 25% premium over base input, charged when the prefix is first stored.

A worked budget for one million calls

Assume a classification or routing agent loop. Each call sends 8,000 input tokens and returns 500 output tokens. Of the input, 6,000 tokens are a fixed prefix: instructions, tool schemas and examples. The other 2,000 vary per call. All numbers are mine, for illustration. Swap in your own.

Scenario (1M calls, 8,000 in, 500 out) Input cost Output cost Total

| Haiku 5.5, no caching | 8,000M tokens x $0.10 = $800 | 500M x $0.50 = $250 | $1,050 |

| Haiku 5.5, 6,000-token prefix cached | 6,000M x $0.01 + 2,000M x $0.10 = $260 | $250 | $510, plus about $7.50 of cache writes |

| Haiku 4.5 ($1.00 in, $5.00 out) | $8,000 | $2,500 | $10,500 |

| Gemini 3.5 Flash ($1.50 in, $9.00 out) | $12,000 | $4,500 | $16,500 |

The cache write line assumes one prefix write per 100 calls, which is 10,000 writes of 6,000 tokens at $0.125 per million. That is a hot-cache assumption. If your traffic is bursty and the five-minute cache expires between bursts, you pay the write more often, and the saving shrinks.

The ratio is the point. Under identical traffic Haiku 5.5 comes out near one tenth of Haiku 4.5 and under one fifteenth of Gemini 3.5 Flash at its listed $1.50 and $9.00. Gemini 3.5 Flash launched on 19 May 2026 with a cached price of $0.15 and batch pricing of $0.75 and $4.50, so its gap narrows with batch jobs but does not close. A newer Gemini 3.8 Flash is reported at an introductory $0.75 and $3.75 through 31 December 2026, then $1.50 and $7.50. Sources disagree on those figures, so check Google's page before relying on them.

OpenAI's GPT-6 Luna is the direct competitor. VentureBeat describes Haiku 5.5 as priced in line with it. I do not have a verified Luna price sheet from this week, so I am not putting a Luna row in the table. Price your own workload on the AI Model Cost Calculator with both models side by side.

Staying under the 100K line

The 1M context window is a capability, not a budget. Using it flips every call into the higher tier. Three habits keep an agent loop in the cheap lane.

First, cap the prompt in code, not in hope. Count tokens before you send. If a request crosses 90,000 tokens, summarise or drop the oldest turns instead of sending it. A ten percent safety margin costs you little and removes the cliff from your daily risk.

Second, treat tool output as the usual culprit. Agents that read files, web pages or search results grow their context fastest through tool results. Truncate results at the tool boundary and return identifiers the agent can fetch on demand, not whole documents.

Third, route the rare big job elsewhere on purpose. If a task genuinely needs 400K tokens of context, it is a different workload. Price it at the higher tier deliberately or give it to a model built for long inputs. Do not let it arrive by accident through a loop that never trims history.

// guard before every call (pseudo-code)
const tokens = countTokens(messages)
if (tokens > 90_000) messages = compact(messages)  // summarise old turns
send(messages)

The new Agent Run Cost Simulator models a multi-step run, not a single call, which is where the tier line bites: step ten of an agent run carries the history of steps one to nine. Run your expected step count through it with the 100K threshold in mind.

What the cheap tier does not fix

Four costs survive the price cut, and they are the ones that surprise teams after the first invoice.

Retries multiply everything. A loop that retries a failed step three times triples the spend on that step, and an agent that wanders can retry silently. Put a hard cap on steps and attempts per run, and log the count per run so you can see the tail. A run that costs 40 times the median is a bug, and it hides inside an average.

Output tokens cost five times input at the lower tier. At $0.50 against $0.10 per million, a chatty model is expensive in a way a terse one is not. In the budget above, output was 24% of the uncached bill but 49% of the cached one. Once you cache the prefix, output is where the money goes, so ask for structured, short answers and set a maximum output length on every call.

Cache misses are silent. The cache read price is one tenth of base input, but only a hit earns it. Change one character early in the prefix, reorder your tool definitions, or insert a timestamp near the top, and the next call pays full price and pays a write on top. Keep the stable part first and the volatile part last, and check the cache hit numbers your provider returns on each response.

Rate limits are a separate ceiling. A million calls a day is about 12 calls a second sustained. Your account limits, not the price, may decide how fast you can run that. Check the limits for your tier before you promise a batch finishes by morning.

None of this argues against the model. It argues for measuring one thing before you scale: cost per successful task, not cost per call. Divide your total spend by the number of tasks that finished correctly. That figure is the one that stays honest when prices, models and prompts change again next month.

Where Haiku 5.5 fits in an agent stack

A model this cheap changes the architecture question from "can I afford a call here" to "which calls do I still send to a bigger model". Haiku 5.5 is a fit for routing and triage, extraction from structured text, log and ticket classification, first-pass code review comments, and the many small steps in a computer-use or browser agent. It is the wrong default for the single hard step that decides whether a run succeeds. For that step, pay for a stronger model and keep Haiku on everything around it.

The honest limit: Anthropic's claim that Haiku 5.5 beats Haiku 4.5 is a vendor claim, and a lower price tells you nothing about your accuracy. Run 200 of your real inputs through both models before you switch, and compare the failures, not the average score. A cheaper model that needs a retry on 15% of calls is not 90% cheaper.

Cost control is the discipline that makes cheap models safe. The AI Agent Ops Bundle covers specs, observability and cost control for exactly this kind of loop, and the Agent Prompt Vault gives you 50 production prompts with stable prefixes you can cache. Fleets of agents also have a coordination bill, which the post The Multi-Agent Tax covers. If your volume is on the OpenAI side, the sibling post on Codex Cloud pricing shows how a rate limit differs from a rate.

Quick answers

How much does Claude Haiku 5.5 cost?

For prompts up to 100K tokens it is $0.10 per million input tokens and $0.50 per million output tokens. Above 100K tokens it is $0.50 input and $2.50 output. Cache reads are $0.01 per million and 5-minute cache writes are $0.125 per million.

What is the Haiku 5.5 context window?

One million tokens, with up to 128K tokens of output. Using more than 100K tokens of prompt moves the call to the higher price tier.

Is Haiku 5.5 really 90% cheaper than Haiku 4.5?

At the lower tier, yes: $0.10 and $0.50 against $1.00 and $5.00. Anthropic's own average figure is 75% cheaper, which is the number to use if some of your calls cross the 100K threshold.

Which models compete with Haiku 5.5 on price?

VentureBeat says OpenAI GPT-6 Luna is priced in line with it. Gemini 3.5 Flash lists at $1.50 input and $9.00 output, so it costs about 15 times more at the same token mix.

Every product mentioned is available at wowhow.cloud — pay once, ship forever.

Originally published at wowhow.cloud

Read the whole story
alvinashcraft
43 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Code Quality Never Changes – Or Is AI Challenging This Notion?

1 Share

This post was adapted from Kai Schmithuesen’s talk at JetBrains Game Development Day 2026. See the full video presentation and Q&A below, or keep reading!

First we are going to define what we mean by code quality then have a look at how this impacts business outcomes, where AI comes in, AI’s impact on quality and how to safeguard it in your team.

You can take the methodologies described in here and use it for many solutions but we will also explore a Qodana example.

Code quality is a topic that tends to get overlooked at times, especially in game development where you might have other issues on your plate. However, speaking with a lot of JetBrains customers at game development conferences, there is an awareness that this is an important topic. That being said there isn’t a set of industry best practices in the gaming industry and so we share some of these best practices from other industries in this presentation.

What do we mean by code quality?

This can be “fuzzily” defined for such an important topic. If you speak with ten colleagues it could be that you get a lot of different answers. Point one, is correctness.

Correctness: Does my code do what I designed it to do?
Performance: Important for gaming, is it fast enough?
Stability: Does it crash or not crash, and how stable it is from a players perspective.
Security: Also important where you don’t want to show up in an industry news story because you got hacked.
Maintainability: Does my code work today and will it work in a couple of years as well?Especially with some games going on for 10 to 12 years.
Reusability: This is a topic getting more relevant for larger studios: can I reuse the code I’ve already developed for other projects? And all of these give you an overview of what code quality entails.

If you go to the wild west that is X or LinkedIn, for example, you can see that whether or not code quality is important is a hotly debated topic. From one person saying they don’t care how their code looks, to the other extreme “My code needs to be beautiful and spotless before I release it,” and anything in between. There are loads of different opinions out there.

The Qodana team advises that code quality can have a real impact on your business and how your game is received. There was a AAA studio that lost 33% of its value a week after launch.

70% was cut from GTA load times by simply fixing two small code flaws. Time to load was cut from 6 minutes to 2 minutes (by a player) and then Dark Souls was offline for 2 months because there were remote code execution flaws (security issue) and if you lose players over these two months they are most likely not coming back and there is real-life impact on revenue. Bad code doesn’t only stay on your side. It lands with your players which can lead to refunds, bad reviews.

When we look at IGN or PC Gamer, quality, bugs, how stable is it, always influences scores. This ultimately determines how many players you will attract and buggy games lead to disgruntled payers. There are some examples of big studios able to get away with it, like Bethesda but this is something you shouldn’t risk.

Why is bad code released?

In a lot of conversations, we find that just the pressure of getting your game to market is immense. Scheduling is a real issue and is reflected in the number of game developers who had to do crunch periods or extended hours – which is over half. This entices a lot of studios to figure out how they can use AI to help with this topic.

AI, especially in game development is a hotly discussed topic. However, 95% of Unity developers (2026 Unity Game Developer Report) use it at work for assistance in the coding process so this shows that AI if it is used correctly can be a helpful tool.

You can prototype much faster, try out new game mechanics, etc. In the past you had to spend a couple of months but now you can use a couple of days and throw out what doesn’t work. You don’t have to spend so much time on boiler issues which results in increased efficiency.

Qodana AI

The goal is not to cut developers from the teams, only enable your team to use AI to do something more interesting. However, this speed and the usage of AI comes with a potential downside when it comes to code quality. AI can generate a lot of code real fast but it’s not necessarily better or worse but does exasperate the amount of issues. In the past you may have written a couple of hundred lines. Now, in the same time AI can write 10,000 lines. This leads to real-life issues for example:

Slow code, is my game up to scratch when it comes to performance, is it stable? Is it exploitable or does it have potential security issues? The good and bad news is that AI code is not necessarily worse than human-generated code, it seems quite on par when it comes to the overall amount of issues but it fails in different ways compared to human-generated code.

When it comes to logic and errors, it creates 1,7x more issues than human-generated code. When it comes back to maintainability etc. is it worse and most concerning it will also potentially create more security issues than a human developer would. This is coming from a study by CodeRabbit. They are an AI vendor themselves.

Qodana AI

However, it’s not necessarily what kind of issues are created by who (humans or AI) but the only question that we pose is “Is our code overall up to scratch and does it meet our quality standards and this is where static analysis can help. The good news is that you can use static analysis to basically cover all of these 6 areas.

Using static code analysis to solve problems for game developers

When it comes to correctness, you have standard inspections. You can also check for code coverage thresholds, which are becoming more important. Not only in coverage but how to use them as well.

Performance is an area that static analysis can cover. Stability a couple of examples would be null safety, resource leak, exception inspections, security inspections and software component analysis – checking for outdated dependencies or doing taint analysis as part of your static analysis to find more complex issues like SQL scripting and so on.

From a maintainability perspective, you can locate code smells, code complexity and duplication. This is something AI is really well-known for. It likes duplications but it doesn’t like refactoring. if you don’t have a way to keep that in check, this will really hurt your maintainability going forward, as no one will be able to understand your code at some stage.

The same with reusability, duplicate code makes it more complicated and static analysis can check if your code adheres to your own internal coding standards, in order to have the same approach – regardless of the project that you’re working on. This way you are also able to share your code much more easily.

AI creates these kind of issues, the most common reason is because it got trained on outdated sources. These are all sucked dry and there isn’t much new work coming up which is what can lead to outdated dependancies and similar.

Static analysis in conjunction with AI as the answer

The downside is, this kind of analysis has been around for a long time but for good reason. One is the results are deterministic.

If you give a static code analysis tool an analysis ten times, it will give you the same set of results ten times. If you run it ten times, it doesn’t cost you more too because it doesn’t consumer any tokens. If you give AI the same piece of code 10 times you’ll get 20 answers.

It’s fast and cheap, if you compare an AI reviewer to static analysis – usually static analysis is much faster. Coming back to being deterministic, you can ties every result to a specific inspection and see why it’s an issue.

This will become more important as compliance keeps growing. One big one is the new EU Cyber Resilience Act already in effect in some areas. Part of this act requires you to create a software bill of materials. If you use AI this will be difficult. If you use static code analysis and you can show a 1-1 relationship between the result and the source you are more likely to comply. However both have their pros and cons.

Pros and cons of static analysis and AI analysis

It can’t understand the intent of your code because it basically looks for patterns. This is where AI comes in. If you want to check that the code does what you want it to do across files or across your project for check for logic-related issues, that’s where AI comes in.

So Ideally, you’d use a combination of static analysis to cover the base and AI on top of that. A human writes the code and AI can also write some . Then once you kick off your pipeline, static analysis covers the initial analysis. Qodana can check for any issues but can also create quick-fixes automatically. So you can still automate even without AI.

These quick fixes are not generated by QIA so this gives you the advantages of static analysis in general, fast, doesn’t cost extra, and reliable. Then you will move on to your AI tool of choice and review the rest and create fixes for other complex issues like a logic issue or refactoring. This can compliment your pipeline.

Ideally you then have a loop, AI will push it back in the pipeline to your static analysis so that static analysis can reconfirm that what AI created is up to scratch and passes your code quality and security standards. Then you can merge the code into your main branch.

It sounds time-consuming but the irony is that this probably takes a couple of minutes each time and can save you time in future. These are a couple of minutes to not spend hours later if there were uncaught issues.

What do the numbers say?

There are already some studies out there that show the advantages of this approach. If you combine static analysis with LLM calls, then the amount of token usage was going down by 72 to 92% while still increasing the code used in that study. The second study found that if you use this combination then you can have up to 33% fewer vulnerabilities in your LLM code going into your master branch.

Here is a quick example of how this can look in Qodana or a static analysis tool of your choice. We support the most common gaming engines out of the box. This is what it looks like in Qodana itself. Normally as a developer you wouldn’t spend much time in here. This is nice for a team lead or a project manager. One of the Qodana team vibe-coded this example project below. It was a banking front-end written in Java and the vibe-coding part itself just took a few minutes to get it up and running.

Watch Qodana Demo

Qodana found 126 issues and gives you a nice way to give you a quick view into what is important. Then you can apply quick-fixes to them. There are different ways to set this up (with varying levels of automation).

Here we’ve done it as part of a GitHub Actions pipeline. Then we put the remaining problem into the Baseline. Once they are in the Baseline they don’t slow you down. The next time you run a scan with Qodana and you find new issues, ideally it would fix 3 automatically, and you’d only see 2 new ones. So it’s easier to figure out how to spend your time.

If we look at the five remaining issues, you can see some show duplicated code. While you can view them from Qodana, you don’t fix them from here. So if we open this in IntelliJ IDEA. You can see it brings us to the exact position in the IntelliJ IDEA codebase.

This will work for our IDEs however, it also works with Visual Studio, Visual Studio Code and Cursor. You can also use our CI tool as well. Here you can review the issues, like duplicated code, by hovering over it (depends on your solution) and then go to more actions, AI actions and have AI fix these issues for you.

You can do it manually or with an agent, semi-automating it so it works in the background and you wait for the clean code on the other end. If you are interested in finding out more, please contact us and we can show you how to use it or you and your team can try it out for free.

Try Qodana

Users have also tried Qodana for game development. Read how Doc Bok used Qodana for a Minecraft game – or how Qodana works on Unity and Unreal Engine projects.

Once you ask AI to add tests, eventually these tests and all repositories start to overflow and even trivial processes like linting can lag. Do you have experience with this?

In some cases you have to jump through ten different loops to release code. Most of the time this is due to someone starting but no one ever reviews the pipeline, what makes sense and what not but static analysis shouldn’t slow you dow.n So, in Qodana, a so-called pull request mode only checks the changed files and should only take a minute or two.

How much can the JetBrains built-in agents in the IDEs read the IDE’s static analysis tools?

Rider has a feature called Hooks. Every time the agent generates code, it’ll call the Hooks in rider. One of them will be to reformat the code and the other to check for the current problems in a file, do the linting, check the problems and report that back to the agent, then the agent can regenerate the code using your standards or fix what you only ask it to expect. It is deterministic so the agent is forced to use the hooks rather than being advised. The hook is enabled and executed every time.

Please note: Widely adopted AI technology is relatively new and the results of burgeoning data should always be deeply interrogated, including data presented in this post.

Special thanks to Kai for his insights.

Read the whole story
alvinashcraft
2 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Famous quotes we've been getting wrong, from Robert Frost to 'Don't mess with Texas,' with Eli Burnstein

1 Share

1227. This week, we talk to Eli Burnstein, author of "Phrased & Confused." We look at why Robert Frost's "road less traveled" wasn't really less traveled, and why the other half of Stewart Brand's "information wants to be free" will make you think about it differently. Then we look at why "Hell is other people" isn't about crowds and how "Don't mess with Texas" began as an anti-littering campaign.


A hearty thank you to the Keepers of the Commas and one Immortal on Patreon. We appreciate your support!

  • Larry Rosenblum
  • Mahala Russell
  • George Wilson
  • Laurel Paul
  • Linda Cox
  • Birna Anna Björnsdóttir
  • Kate Wade


🔗 Share your familect recording in Speakpipe or by leaving a voicemail at 833-214-GIRL (833-214-4475)

🔗 Watch my LinkedIn Learning writing courses.

🔗 Subscribe to the newsletter.

🔗 Find an edited transcript.

🔗 Get Grammar Girl books.


| HOST: Mignon Fogarty

| Grammar Girl is part of the Quick and Dirty Tips podcast network.

  • Audio Engineer: Dan Feierabend
  • Director of Podcast: Holly Hutchings
  • Advertising Operations Specialist: Morgan Christianson
  • Marketing and Video: Nat Hoopes, Rebekah Sebastian
  • Podcast Associate: Maram Elnagheeb

| Theme music by Catherine Rannus.

| Grammar Girl Social Media: YouTube. TikTok. Facebook. Threads. Instagram. LinkedIn. Mastodon. Bluesky.


Hosted on Acast. See acast.com/privacy for more information.





Download audio: https://sphinx.acast.com/p/open/s/69c1476c007cdcf83fc0964b/e/6ac037295558f9dbe286a0a7/media.mp3
Read the whole story
alvinashcraft
2 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Narrative intelligence and the human advantage

1 Share

As AI reshapes how we work, create, and think about our roles, staying adaptable means understanding not just the technology, but our own stories and identities. Reed Frerichs joins Daniel and Chris to explore narrative intelligence, storytelling, leadership, and personal growth in an AI-driven world. They discuss how entrepreneurs and leaders can use AI tools to navigate changing roles, embrace uncertainty, and develop emotional intelligence along with Reed's “own, author, act” framework.

Featuring: 

Links:

Sponsors:

Resources and Events:





Download audio: https://pscrb.fm/rss/p/dts.podtrac.com/redirect.mp3/media.transistor.fm/f2fcb677/ee0419b3.mp3
Read the whole story
alvinashcraft
2 hours ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories