Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
160455 stories
·
33 followers

Anthropic’s Fable 5.1 is a bit cheaper, a bit smarter, and refuses a lot less

1 Share

On Tuesday, Anthropic launched the latest versions of its flagship Fable and Mythos models. Anthropic promises that the updated models will offer stronger performance at a lower cost, largely because the company is reducing cache read pricing.

Like before, Fable 5.1 and Mythos 5.1 are the same model under the hood, but while Fable 5.1 is now generally available, Mythos 5.1 remains restricted to Anthropic’s trusted access program because it features safeguards that Anthropic says are “specifically designed to support work in cybersecurity and the life sciences.”

Pricing remains unchanged at $10/$50 per million input/output tokens, but cache reads are now only $0.25 per million, down 75%.

Fable 5.1, like Fable 5, remains restricted to API usage and users with Max, Team Premium, and Enterprise Premium plans (at 50% of their weekly usage limits). Pro subscribers can use it via usage credits.

Fable 5.1 benchmark results

In the benchmarks Anthropic shared, Fable 5.1 easily outperforms the older model, Anthropic’s Opus 5, and OpenAI’s GPT-5.6 Sol. Nobody is bothering with comparisons to Gemini 3.1 Pro anymore.

In general, Anthropic says the new model outperforms Fable 5 in virtually all aspects at a lower cost. This depends a bit on the benchmark, but for the most part, it comes down to being able to step down the reasoning mode by one or two notches and still get the same results as Fable 5 at higher settings, all with Fable 5.1 using fewer tokens overall.

In general, Anthropic says the new model outperforms Fable 5 in virtually all aspects at a lower cost.

Many of the improvements aren’t dramatic, except for a major step-up in Fable 5.1’s performance on the Terminal-Bench-Science 0.1 benchmarks, which test the model’s capabilities around using agents for scientific research. Here, Fable 5.1 scored 52.6% versus Fable 5’s 24.7%.

In most other areas, the improvements are within 2-4% of either Fable 5 or Opus 5.

Because of this, Fable 5.1 in Claude Code defaults to the high reasoning mode setting, but medium on Claude Cowork and Claude.ai.

Claude Code safeguards loosen up

One issue with Fable 5 was that it was often too cautious and deferred to the Opus models when it sensed that the user was pushing its safeguards. Now, Anthropic says, it has made these safeguards more precise, “ensuring that they’re less likely to flag benign content (like queries about medical issues or cyber defenders using the model to make their systems safer), but still ensuring they provide robust protection against genuine threats.”

For coders who use Claude Code, the most immediate change is that the model’s safeguards will hopefully get out of the way more often.

For coders who use Claude Code, the most immediate change is that the model’s safeguards will hopefully get out of the way more often.

Anthropic says its cyber safeguards now intervene about 60% less per session than they did with Fable 5, and the model is now allowed to identify vulnerabilities in code.

Penetration testing, exploit generation, and scanning binaries for vulnerabilities remain off the table, though, and Anthropic says it will continue to redirect those requests to its Opus models.

The biology safeguards got a similar update and now intervene 85% less often on benign requests about basic biology and medical questions.

What about those watermarks?

Anthropic recently shared its technique for watermarking text generated by its models (for models released after August 2, 2026.

The watermark is an invisible numerical marker that signals how likely it is that Claude wrote a given piece of text. Anthropic says it doesn’t affect output quality and contains no data about the user, organization, or conversation. Until now, though, nobody outside Anthropic had a way actually to check for it.

With this release, Anthropic is opening a detection API in private preview “to eligible organizations as required under EU law (such as regulators, law enforcement, media, fact-checkers, independent researchers, educational organizations, and EU civil society groups).”

Enterprises that are obligated to verify watermarking to comply with the EU’s AI Act will also get access, with a wider rollout planned for the future.

Data retention

One point of contention for enterprises that wanted to use Fable 5 was Anthropic’s retention policy. While Anthropic offers zero data retention to qualifying customers on its other models, Fable 5 users had to opt into a 30-day retention period.

Now, Anthropic is rolling out what it calls Enterprise Frontier Safeguards (EFS), which it says delivers the privacy of zero data retention while preserving its ability to monitor for misuse.

“EFS works by storing data in cloud infrastructure controlled entirely by the customer, not Anthropic,” the company writes in its announcement. “It will be made available to enterprise customers in phases, beginning later this fall. Until EFS is available, eligible customers will be able to use Fable 5.1 with zero data retention.”

Anthropic says it designed this new system based on customer feedback. Businesses in regulated industries, after all, were essentially unable to use the model under those data retention rules.

“We therefore sat down with customers to design a solution that could provide the best of both worlds: the privacy of ZDR and the safety allowed by monitoring across time and accounts,” Anthropic explains.

In practice, this means EFS moves that data into the customer’s own cloud storage, such as Amazon S3, Azure Blob Storage, or Google Cloud Storage, where it remains encrypted with the user’s encryption keys. It then runs Anthropic’s automated detection against it with no human review on Anthropic’s side, and routes any alerts back to the customer.

As of now, this tooling covers Claude Code, Claude Enterprise, the Claude Platform, and Claude on Amazon Bedrock, Microsoft Foundry, and Google’s Agent Platform. The company says there will be no additional cost beyond the customer’s own cloud storage bill.

EFS will roll out in phases starting this fall. Until then, Anthropic will not retain any data from usage of Fable 5 and Fable 5.1.

Anti-distillation

Fable 5.1 also ships with what Anthropic calls strengthened mechanisms to avoid distillation. The company, of course, has been quite vocal about its suspicions that Chinese labs have been distilling its models at an industrial scale.

The first of them is a change to the API itself: New accounts can no longer “manually edit Claude’s prior context in a multi-turn conversation while preserving the transcript of Claude’s prior thinking.”

This, Anthropic says, closes off “a common, publicly documented distillation technique.” However, it’ll also trip up agent harnesses that rewrite or compact conversation history, and the company says the restriction will extend to existing accounts with future model releases.

The post Anthropic’s Fable 5.1 is a bit cheaper, a bit smarter, and refuses a lot less appeared first on The New Stack.

Read the whole story
alvinashcraft
33 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities

1 Share
The company will give select partners early access to its Astra AI model—so they have time to shore up their defenses.
Read the whole story
alvinashcraft
34 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Enterprise Architecture in the AI Era: Tools, Capabilities, and the Road to Autonomy

1 Share

An enterprise architecture (EA) tool is a software platform that enterprises use to capture, connect, and continuously maintain a structured picture of the enterprise covering strategies, business capabilities, processes, applications, data, technologies, and the relationships between all these elements.

An EA tool acts as a Central Enterprise Repository (a “single source of truth”) that architects and other stakeholders use to model both the current state of the enterprise and the desired future state.

Read the whole story
alvinashcraft
34 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Fragments: September 1

1 Share

Like many readers, I’m wary of AI generated prose. Simon Wilison has written an LLM cliché highlighter - paste in some text, or a URL, and it will flag various patterns common to LLMs. It references a wikipedia page of signs of AI writing. That page points out that:

Humans are notoriously bad at distinguishing human and LLM-generated text. While research on humans’ abilities to detect AI-generated text is still limited, a 2025 study has shown that human ability to distinguish LLM text from human is no better than random chance. Another 2025 study on German theses has shown that humans managed a “recognition rate of 57% for AI texts and 64% for human-generated texts”.[

Not just do I find myself repelled by prose with an LLM-voice, I also wonder how accurate my reaction is. I’m old enough to see all sorts of new tic-phrases appear, and in the past would just chalk it up to youngsters or airport business books. (Not to mention Americanisms, which I’ll get used to momentarily.)

 ❄                ❄                ❄                ❄                ❄

NVIDIA’s technical blog reports on an Architecture for Long-Horizon Autonomous Agents. Their research group used a combination of Claude Opus 5 and a harness called AVO, and used it first to do GPU kernel optimization and then a broader reasoning benchmark (ARC-AGI-3). Both of these were long-term tasks, for the kernel optimization the agent ran for seven days.

AVO is designed to preserve progress beyond a single model context. Two mechanisms are particularly important: persistent memory and supervision.

Persistent memory carries forward prior implementations, evaluation results, compiler and profiler outputs, and accumulated reasoning, allowing the agent to resume from the current state rather than repeatedly reconstructing the search.

The supervisor monitors the broader trajectory for stagnation or repeated unproductive cycles and can redirect the main agent toward alternative strategies when needed. During the seven-day attention-kernel run, the main agent remained responsible for deciding what to inspect, change, test, and evaluate, while the supervisor helped maintain forward progress when the search plateaued.

The team was encouraged that AVO did well at two different kinds of long-horizon tasks, indicating that it’s a general-purpose tool.

 ❄                ❄                ❄                ❄                ❄

Mickey Petersen:

MCP is SOAP for Zoomers.

 ❄                ❄                ❄                ❄                ❄

Paul Stack writes that AI Broke the Assumptions Behind CI. Here’s his description of CI with agents.

An agent writes a change, opens a PR, and CI picks it up instantly. The compile fails, the agent pushes a fix, CI picks it up instantly again. A test fails, another fix, another instant run. Each iteration is fast, but the agent is still discovering that its change doesn’t work only after it crosses the PR boundary. The feedback loop is in the wrong place regardless of how fast CI runs.

He points out that all of this breaks the pipeline, because “CI” keeps failing, and advocates doing verification before the agent pushes. This is where I get to be the grumpy old guy, and point out that was always how Continuous Integration works. When I’m done with a change, first I pull (to get everyone else’s change since I started), I build and test locally, and if all is well I push and let the CI server do its thing. The only reason the CI server should fail is if there’s some funky mismatch between my machine and the CI server. Tests that take a while to run aren’t part of this loop, instead they are run further down the deployment pipeline, downstream of CI. Any failures there imply missing tests in CI.

(I’m being a bit unfair dumping on this article here. After all I could have filled a full working day correcting misleading descriptions of Continuous Integration for most of the last twenty years. Maybe I’m just after an excuse to point readers to the extensive range of articles hosted here about what’s needed to get code from laptop to production.)

Stack is right that we should question how the deployment pipelines should work with agents in play. He’s also right that CI with humans relies on them being disciplined to run commit tests locally before pushing to the CI server - and that we can (and should) automate that when using agents. I also don’t know more about his setup than what he’s written in his post, so there’s likely complications he faces that I don’t understand. But when thinking about designing pipelines it’s important to understand the principles that underlie Continuous Delivery, understand how the practices really work, and understand why they are in place. Above all, Continuous Integration is a practice, not just the CI server. Yes, CI does conflate two jobs: executing verification and coordinating merges. But that’s the point: verification is a necessary part of merging if we want to retain a healthy mainline.

 ❄                ❄                ❄                ❄                ❄

Recently Noah Smith posted an article about how he was worried about an AI-generated super-virus savaging humanity. It’s a worry I’ve heard a few times, seen as a greater concern than AI turning us into labradors or paper-clips. Claus Wilke, who works in the field, isn’t so concerned.

Computational design of biological systems is unfathomably difficult. Experts who have dedicated their life to this topic routinely hit their head against the wall when nothing they try seems to work. PhD students in 2026 using state-of-the-art AI software are spending months or years trying to design simple peptide binders that inhibit some enzyme or pull down some protein, and the majority of their designs fail, or don’t express, or are toxic. But in Smith’s fictitious world a disgruntled teenager with no special training in biology can just solve a problem thousands of times more complicated than designing a peptide binder. The distance between where we are today and where we would have to be for Smith’s story to have any realism is enormous.

 ❄                ❄                ❄                ❄                ❄

It seems that a couple of remarkably talented academics are experts in a staggeringly wide range of fields. Or maybe they are just ghosts.

These names do not exist. Elena Vasquez and Marcus Chen have appeared as volcano experts, astronauts, thriller protagonists, podcast hosts, and academic co-authors across hundreds of independently produced AI-generated documents, never having lived. We show that large language models do not merely default to high-probability individual names when generating fictional experts: they produce correlated character ensembles: pairs and trios whose co-occurrence rates far exceed chance and are consistent across independent generations.

Read the whole story
alvinashcraft
35 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Build Your Personal Brand with GitHub Copilot — Full Course for Students

1 Share
From: Microsoft Developer
Duration: 2:05:53
Views: 16

Go from zero to a live portfolio website using the GitHub Copilot app — no coding
experience required. This is the complete 10-episode student series in one video,
covering the Student Developer Pack, prompting, Git and GitHub Pages, context and
tokens, MCP servers, skills and agent personas, bring-your-own-model, Bootstrap,
and controlling Copilot from your phone.

In this course, you'll learn:
→ How to claim the GitHub Student Developer Pack and install the Copilot app + VS Code
→ How to turn a resume PDF into a real portfolio website by prompting, not coding
→ What Git, commits, branches, diffs, and pull requests actually are — explained on a lightboard
→ How to publish your site free on GitHub Pages and redeploy every time you edit
→ How to control Copilot with Interactive, Plan, and Autopilot modes
→ How context windows and tokens work, and how to use /init, /context, and /compact
→ How to extend Copilot with built-in tools and MCP servers like Playwright
→ How to build a reusable agent skill and install skills from a plugin marketplace
→ How to bring your own model with Microsoft Foundry and the Azure for Students credit
→ How to run a Copilot session from your phone with remote control and a QR code

CHAPTERS
0:00 Welcome — what you'll build in this series
1:05 Claiming the GitHub Student Developer Pack
7:15 Your first prompt: a Hello World site
10:45 EPISODE 2 — Build a portfolio website from a resume
14:23 Answering Copilot's questions about structure and style
18:37 Resume preview, school colors, and dark mode
23:05 EPISODE 3 — Git, GitHub, and GitHub Pages
23:42 Version control explained: diffs, commits, branches, merging
32:56 Creating a repository and pushing your site
47:49 Deploying with GitHub Pages
54:11 EPISODE 4 — Copilot settings, sessions, and modes
56:23 Interactive mode, Plan mode, and Autopilot mode
59:33 EPISODE 5 — Context and tokens
1:02:53 Adding project files and running /init
1:05:52 Shrinking your session with /compact
1:06:40 EPISODE 6 — Tools and MCP servers
1:11:08 Installing and using the Playwright MCP server
1:15:33 EPISODE 7 — Skills, agent personas, and marketplaces
1:23:50 Turning a workflow into a reusable skill
1:30:22 Adding a plugin marketplace
1:37:02 EPISODE 8 — Bring your own model
1:40:43 Using the Azure MCP to create a Foundry resource
1:47:48 Connecting Foundry as a model provider
1:51:50 EPISODE 9 — Modernize the site with Bootstrap
1:54:05 Sticky navbar, animated cards, and shadows
1:59:31 EPISODE 10 — Control Copilot from your phone
2:01:51 Remote control and the QR code handoff
2:04:29 Back at the laptop — the finished feature

https://aka.ms/StudentAI-GitHubCopilotApp
https://aka.ms/StudentAI-GitHub
https://aka.ms/StudentAI-GitHubPack
https://aka.ms/StudentAI-VSCode
https://aka.ms/build-your-personal-brand-with-copilot
https://aka.ms/student-learning-series
https://aka.ms/student-demo-series-website

Read the whole story
alvinashcraft
1 hour ago
reply
Pennsylvania, USA
Share this story
Delete

Quadratic Regression with QR-Householder OLS Solve Training Using C#

1 Share
Dr. James McCaffrey demonstrates how to implement quadratic regression from scratch in C# using QR-Householder OLS solve training with L2 regularization, covering the underlying matrix operations, model training, evaluation and prediction.
Read the whole story
alvinashcraft
1 hour ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories