Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
159540 stories
·
33 followers

Claude Sonnet 5.5: A Game-Changer in AI Pricing and Performance?

1 Share

The landscape of artificial intelligence is ever-evolving, constantly challenged by the release of cutting-edge models that push the boundaries of what machines can comprehend and execute. One recent development that’s been stirring the buzz is the potential emergence of Claude Sonnet 5.5, codenamed “Fennec,” reportedly being developed by Anthropic.

For those unfamiliar, context windows in AI models determine how much information the AI can juggle at once. Rumor has it that this new model might possess a staggering 2 million token context window, a significant leap from current standards, offering faster inference, enhanced tool usage, and potentially a price that remains competitive within the Sonnet tier.

Based on content from BitBiasedAI

Unpacking the Leaked Information

The excitement arises from a report indicating Claude Sonnet 5.5 could eclipse the current market players, not only in capability but crucially, in pricing. Yet, it’s vital to approach this news cautiously, as it stems from a single-source leak, with no official confirmation from Anthropic.

The code name “Fennec” isn’t new to the AI community; it’s been previously linked to other rumored Sonnet releases, which either did not materialize or differed from initial expectations. This recycling of codenames highlights the unpredictable nature of insider reports, reminding us to temper our anticipation with skepticism until more information from credible sources is available.

The Implications of a 2 Million Token Context Window

What would a 2 million token context window mean in practice? Such an expansion could allow the model to process significantly larger datasets in a single pass, whether that’s a gigantic codebase, exhaustive legal documents, or extended transcripts. This increased capacity is particularly beneficial for agentic workflows where multiple tool calls are integrated, ensuring continuous context knowledge without the risk of data truncation.

However, it’s worth noting that simply having a larger context window isn’t a surefire solution. The effectiveness of such a model depends on its ability to maintain accuracy as it delves deeper into this extended context, a phenomenon known as “context rot.”

Pricing Versus Performance: The Core Question

The central narrative is less about the sheer size of the context window and more about the value equation Anthropic might offer. If they can deliver this expansive capability while keeping the cost below other flagship models, Claude Sonnet 5.5 could disrupt the industry’s pricing dynamics significantly.

Imagine a model offering near-flagship performance at a mid-tier price, drastically altering which tools developers and businesses opt for based on cost-effectiveness instead of absolute capability.

The Bigger Picture for AI Enthusiasts

For those currently leveraging AI models, this development poses a strategic dilemma. Should one hold out for unconfirmed advancements or proceed with the available options like Opus 5 or Fable 5? At present, the prudent course of action would be to continue with existing models while keeping an ear to the ground for confirmed developments regarding Sonnet 5.5.

Analyzing the larger theme—how Anthropic seems set on closing gaps in capability while maintaining competitive pricing—is where enthusiasts and experts should focus. This strategy, proven through past releases, suggests a tangible shift in AI affordability and accessibility on the horizon.

Conclusion

In the AI domain, speculative narratives can often overshadow substantiated trends. While the Fennec leak does support Anthropic’s history of advancing model capabilities without inflating costs, the specifics remain to be seen. For now, the rumored features of Claude Sonnet 5.5 provide a fascinating glimpse into potential future developments—a compelling reason for both excitement and caution among AI stakeholders. Keep watching this space for any further confirmations or additional insights.

For the latest in AI developments, stay tuned and share your thoughts on whether a model like Sonnet 5.5 could redefine your choice of tools. Does context window size hold the sway over your model selection, or does cost-effectiveness define your preferences? Let us know in the comments!

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Fragments: August 18

1 Share

Part of the reason why I’m at Thoughtworks is because I’d like to see a software development organization founded on technical excellence as an example for the rest of the industry. The trouble is that I have little aptitude or inclination for the hard work of building such an organization. So I rely on working with people who are prepared to actually put the effort in. A key partner in all of this is Rachel Laycock, who is the global CTO of Thoughtworks.

Not just is she far better than me at running a technology organization, she’s also a keen observer and connector of ideas. I’ve been urging her to write these down, even if her busy schedule makes it difficult for her to compose them into something substantial.

Happily she’s starting writing “Rachel’s Ramblings”

Fast, imperfect, thinking out loud. Naming ideas early rather than waiting until they’re fully formed. Because the reality is, most of what I do day to day isn’t answering known questions. It’s spotting patterns and asking questions we haven’t quite figured out yet.

 ❄                ❄                ❄                ❄                ❄

My colleagues in Europe are organizing XConf Europe in London on September 11th.

The sessions examine what happens when agentic systems meet compliance, how to run sovereign models, performance patterns in data migrations and how to safely navigate legacy codebases. Lu Wilson will give a keynote on ‘Jam-oriented programming’.

 ❄                ❄                ❄                ❄                ❄

Noah Smith recognizes the high usage of AI, and its impressive feats - but also that there aren’t signs of massive productivity growth or job losses. This may be the calm before the storm, but Smith thinks there may something else in play. He quotes a metaphor from François Chollet

One of the biggest misconceptions people have about intelligence is seeing it as some kind of unbounded scalar stat, like height. “Future AI will have 10,000 IQ”, that sort of thing. Intelligence is a conversion ratio, with an optimality bound. Increasing intelligence is not so much like “making the tower taller”, it’s more like “making the ball rounder”. At some point it’s already pretty damn spherical and any improvement is marginal.

The thought here is that intelligence in the sense that we know it, isn’t something where there’s a lot of room for massive improvement. That doesn’t mean AI won’t be “smarter” than us in other respects, after all even without AI my computer is better at me than remembering what I’ve agreed to do over the next six months.

But even if AI doesn’t get smarter than humans, it can gain by being more replicable. Not just does this make it cheaper to use, perhaps more importantly it makes it more responsive. While I might harrumph at how slowly The Genie responds to my queries, it’s still far faster than contacting a human.

Smith continues by surmising that AI may be able to make sense of phenomena that can’t be reduced to simple laws, but can only be understood by something able to comprehend a multitude of details:

there may be laws of the universe that humans can’t understand but AI can. I call these “cloud laws” — causal regularities that can be exploited by technology, but which are too diffuse and complex for an individual human being to either intuit or communicate.

His thought is that even if there isn’t any space for AI to get more intelligent than humans along the lines we are used to, that they can open up new directions. As well as these cloud laws he also thinks that AI can understand human systems that rely on the kind of tacit, distributed knowledge that human organizations build up over time.

My take-away here is that AI won’t seem more intelligent in the way that we typically frame intelligent, but more intelligent in different ways. The converse of which is that the human value comes in artfully combining our human nature with these new spells that The Genie can cast.

 ❄                ❄                ❄                ❄                ❄

Especially in our profession, we’ve seen increasing emphasis on the importance of data. However I’ve observed that most people still struggle to understand the message data is telling us. One of the reasons I’m interested in election forecasting is in how they communicate their insights, especially since so many people have difficulty with probabilistic forecasts. (I often wonder how much being a board-gamer has helped me be comfortable with this, all that time interacting with Combat Results Tables in my youth must have benefited me somehow.)

50+1 (one of the successors of 538) have published a little explainer on how they designed their 2026 election forecast page. There’s a good discussion of the logic behind their simulation histogram, I like how they use a text annotation to explain one point, giving the reader enough guidance to understand the rest of the graphic. They also tackle the knotty problem of visualizing geographical data on the house races. There’s a common visualization error in the U.S. using choropleth maps that leads to large areas of the landmass shown red, implying dirt votes rather than humans. Their approach to this, using dots on the map, helps visualize both the politics and the population density. They also explain how to deal with this kind of data on small screens. Lastly they describe their approach to tabular data, and how this is the right place for lots of details, together with affordances to help both casual and power-users navigate those tables.

 ❄                ❄                ❄                ❄                ❄

I’ve kept an eye on Alex Stamos for a while now, as he’s a sensible voice on security and safety. He’s posted a newsletter on substack that casts an intelligent eye over recent safety issues with AI.

He makes a clear critique of recent US government actions around LLM models

On a Friday afternoon at around 5pm PT, Anthropic was forced to shut down a system that had been plumbed into coding agents, SOCs, customer service bots, and countless products. […]

This had the immediate effect of injecting political risk into the US AI ecosystem for both American and non-American customers. It signaled that you cannot depend on American AI infrastructure because, at any moment, an unwritten, capricious, and legally dubious justification could be used to yank that infrastructure from underneath your feet.

When Fable was turned back on, it was much dumber and less useful to cyber defenders

[…]

While Fable was down, Z.ai was taking advantage of the free market and permissionless innovation culture provided by the (checks notes) General Secretary, Politburo, and Communist Party of the People’s Republic of China, and released GLM 5.2. With 753B parameters, it falls a bit short of Opus 4.8 in most tasks but is extremely efficient and is small enough to be trained and hosted in many enterprise contexts. With an MIT license it can be fine-tuned with a wide range of techniques and used by any customer in any context. Since then, Kimi K3 has rocked the industry by providing Fable-like performance

As he highlights, one of the biggest dangers with the danger of shutting down a frontier model is that it can cripple an organization’s defenses:

Hugging Face tried to use an Anthropic model to defend itself during an active incident, got blocked by the classifier, and moved to GLM 5.2 on an emergency basis. Their advice to everyone else was to keep an open-weight model on the shelf for defensive cyber.

On the whole, he sees it as a Good Thing that these model escapes have happened:

The OpenAI attack against Hugging Face, and Hugging Face’s excellent write-up has given us a preview of what a standard AI-enabled attack might look like in a matter of months.

It’s good that we got this warning shot. Nobody got hurt, the target was a sophisticated actor with the ability to defend themselves and the ability to give us a detailed write-up, and OpenAI turned the model off.

He follows up by saying that all of this is signal that we should “stop talking about AI finding bugs, focus on fixing them”. These modern LLMs can do much to fix bugs and improve security, and people need to work on that rapidly to fix holes before less reputable folks than OpenAI find them. Then figure out how to harness LLMs to introduce this kind of checking into the everyday build process, so that this kind of analysis just a step in the continuous delivery build pipeline.

I agree with him both that open-weight models should be legal, have their upsides, but will also be used for many bad things by bad actors. Both the industry and government agencies need put serious effort into figuring out how to mitigate these risks.

Where I would go further is to say the same is true of the closed-weight models too. Although closed weight models are subject to greater controls, the same fundamental issues apply. He rightly takes the foundation model companies to task:

There is an old saying I pass down to my students when I give them career advice - if you are a jerk to people on your way up, don’t expect them to catch you when you are on your way down

There’s a lot of sound advice for model companies, the government, defenders, and venture capitalists.

We will go through some rough changes, I just hope that we will indeed come through it with a better society. On the whole, that’s happened with previous technological changes like this, but past performance does not guarantee future results.

 ❄                ❄                ❄                ❄                ❄

The Economist has a good article on the impact of AI in China.

China has made an all-out push in ai, under the conviction that, in its competition with America and the rest of the world, dominance of the technology is an almost existential necessity. […] But the party is increasingly concerned about how ai will displace workers.

Robots and AI are appearing in an economy that’s struggling after the recent property crisis. The Chinese government is opposing firms using AI to cut jobs. China will need robots: its population will shrink by 25% by 2050. But with less working people, there’s less financial support for pensions. Many countries have to deal with shrinking population, but China’s challenge is particularly acute.

 ❄                ❄                ❄                ❄                ❄

Rob Bowley:

I go on holiday for a few weeks and we’ve already moved on from Loop Engineering to Graph Engineering

The half-life of a paradigm is getting shorter than my annual leave

My prediction: neuro-symbolic engineering by the end of August, at which point we’ll have gone full circle and reinvented Prolog

Read the whole story
alvinashcraft
15 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Want to use AI agents safely? Start with design by Colin Eberhardt

1 Share

Time and again, the biggest concern we hear from clients about agentic AI is a loss of control.

This is by no means irrational. AI is non-deterministic, meaning that the same input does not always result in the same output. In addition, AI agents are increasingly being granted full autonomy in making decisions and taking actions, where this non-determinism can have a significant impact. Organisations are right to ask how they can maintain control, accountability and trust.

The risks and concerns that need to be addressed are wide-ranging, including:

  • Quality – will AI agents deliver correct results? Will they be unbiased?

  • Security – will the agents be susceptible to deliberate attack or accidental failures?

  • Auditability – how can you demonstrate that your system is behaving as it’s expected to behave?

  • Operational Resilience – how do you manage the reality that the behaviour of AI systems will change over time?

However, the good news is that organisations have decades of know-how and experience in managing risk in complex systems and dealing with non-deterministic agents: humans. While agentic AI introduces new challenges, it doesn’t overturn decades of thinking about process design, controls and governance.

In this blog, we’ll explain the fundamental importance of design when it comes to introducing AI into your systems and processes. We’ll show how making the right decisions at the design stage ensures that you introduce guardrails that are enablers rather than constraints. And we’ll explain how the design decisions you make will determine the right approach to observability and monitoring, and how those controls can evolve as the system matures. But first, let’s begin with a mistake many organisations make.

Don’t start with guardrails…

At Scott Logic, we have more than two decades of experience in working with organisations in highly regulated environments. It’s therefore unsurprising that the starting point for many such organisations is to ask, “What controls do we need?” The problem with this mindset is that it can constrain the art of the possible. When it comes to leveraging AI, the overzealous application of guardrails can squeeze out any value that might have been gained.

Instead, in our experience, the best Chief Information Security Officers (CISOs) always start by asking “What are we trying to achieve?” and want to understand the end-to-end process. They then ask what controls are required and how their effectiveness will be evidenced.

To leverage AI safely, you should take the same design-first approach: it will help you choose the right technology for the task, determine the appropriate level of autonomy, and decide where human oversight and accountability should remain.

…start with design

In asserting the importance of the design stage, we’re not advocating big, upfront waterfall design. We’re simply recommending that you should spend time at the start to address some fundamental questions and concerns. You should think carefully about the end-to-end process in which the AI agents will operate. Agentic AI creates an opportunity not simply to automate existing ways of working, but to redesign them. The aim should be to identify where humans and AI each create the most value, then design the process around those strengths.

When it comes to the system itself, the same fundamental questions about objectives, risk, control and accountability still need to be answered. As with any project, you need a clear business objective to serve as the North Star that guides your decisions. Much also remains the same in terms of system design. All the design patterns that have evolved over the years to create secure systems still apply, e.g., Identity Management, Principle of Least Privilege, etc.; they just need to be adapted for a system leveraging non-deterministic AI agents.

In the same way, other design considerations depend on familiar, broad principles, as follows.

Autonomy should be proportionate to risk

An early question to answer up-front is what level of risk is involved in the outcome of the end-to-end process that’s being designed. That’s because the higher the risk, the less you should rely on non-deterministic decision-making.

For example, let’s consider mortgages. If you were designing a system to explain mortgage-related concepts in plain English to first-time buyers, the level of risk would be low, and therefore the use of Generative AI (GenAI) would be a sound choice. However, if you were designing a system that would task an AI model with making mortgage-lending decisions autonomously, the use of GenAI would be a poor choice with potentially life-changing and business-damaging consequences.

Accountability remains human

This is a natural extension of the autonomy question. AI cannot be held accountable in law; only organisations and people can. So, the higher the level of risk involved in the outcome, the more human oversight and decision-making is required. Guardrails and observability have no bearing on this; just because you can provide evidence of AI decision-making and behaviour, that still won’t make it accountable.

Use the right tool for the job

Organisations have been using AI for decades, but it was hard work. Lots of time and effort went into ensuring that the AI delivered the desired results. What’s changed is how easy it now is to use AI, and the confidence with which it asserts that it’s provided the right answer. The ease disguises the underlying complexity and the chance of error.

As a result, organisations feel encouraged to apply GenAI to the wrong use cases. There are many different kinds of AI, and this should be remembered at the design stage. For example, algorithmic trading already uses AI, but it is intrinsically different from GenAI. If algorithmic trading harnessed GenAI, the results would likely be disastrous.

Design the end-to-end process

It’s only by mapping out the end-to-end process that you can make the right decisions about how you will maintain control and ensure that the system operates safely. It’s also worth challenging the assumption that existing controls, approvals and handoffs should remain unchanged. Many business processes were designed around the limitations of human decision-making and human effort. It’s by identifying where AI will create value and where humans need to stay in the loop that you will be able to design the right guardrails and observability mechanisms.

Guardrails enable safe autonomy

Guardrails encompass a range of mechanisms applied to a system to define acceptable behaviour, reduce risk and keep operations within agreed boundaries. They may take the form of technical controls, procedural checks, or other measures appropriate to the level of autonomy and risk involved. Their ultimate purpose is to maximise the value a system can create while keeping it safe and controllable.

Deterministic systems have always been designed with deterministic guardrails, using black-and-white rules and controls. What’s changed in the era of non-deterministic AI is that guardrails need to be more adaptable and qualitative in nature, favouring positive non-determinism and reining in negative non-determinism.

Guardrails can go a long way in exercising control over AI. However, it’s important to say that the more you use non-deterministic systems, the more you have to accept that the probability of failure is higher than with a deterministic system. This must be weighed against the value that you’re gaining by harnessing the non-deterministic AI, and also against the probability of failure with previous, human-centred processes.

Due to the higher probability of failure, the role played by observability is more important than before.

Observability helps you understand what’s happening

Observability is the capability of inspecting, understanding and monitoring a live system. Whereas traditional system logging is passive, observability is active.

AI models and systems ‘drift’, which means that their behaviour can alter over time due to numerous factors. Among other reasons, there might be changes in the data consumed by the model, or the behaviour of its users, or the prompts and instructions it receives. If you’re using a third-party, API-based model, the provider may update it or replace it. Meanwhile, you need to keep a close eye on the cost of running the system and protect it against attack; as systems become more dynamic, so too do the risks and vulnerabilities they face.

This makes it critically important to monitor, understand and manage the behaviour of the evolving system throughout its lifecycle. For this, you need a mechanism to evaluate the quality of the system’s outputs. This could include a test suite that helps you continuously monitor behaviours, outputs and outcomes, so that you can refine the model’s instructions. As the system evolves over time, so will the test suite.

In all contexts, but especially in highly regulated environments, observability allows you to demonstrate that your system is behaving as it was intended to do. That means thinking carefully at the design stage about what needs to be observed. Key decision and control points should be identified in advance, including the choices an agent will be instructed to make, the permissions it will exercise, the tools it might invoke, and the level of autonomy under which it will operate.

Those control points are not limited to actions and decisions. Organisations may also need visibility of the identity and authority of both human and AI participants within a process. It’s not enough to know that a particular action took place. Organisations may also need to know which agent performed it, which model or version it was using, whether it was acting autonomously or on behalf of a human user, and what authority it had been granted. This provenance will become increasingly important as agentic systems take on greater autonomy. It’s what will help organisations establish accountability, audit decisions, and detect misuse.

It all comes back to design

We’ve been brought up in a society where computers are integral to our everyday lives and always assumed to be right. To get the most out of GenAI and large language models, we need to accept and accommodate the idea that they will sometimes give us incorrect answers. If we try to constrain AI so that it’s always correct, we will lose most of the value that this technology can deliver. Again, this means that organisations will need to understand that it’s not always the right tool for the job.

Working out where it is the right tool for the job is not easy, especially if you simply try to integrate it into existing processes. As we said earlier, the much better approach is to take a step back and consider how you might reimagine the end-to-end process. At Scott Logic, we have worked with clients like Yuki and Scopevisio to reimagine software engineering in this way with extraordinary results, resulting in a fundamentally different process design. There’s the potential to achieve similar gains by reimagining other business processes, optimising them for a combination of AI and human agents. It’s a design challenge – which Scott Logic can assist you with – and it’s where the real opportunity lies.

Read the whole story
alvinashcraft
19 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers

1 Share
Read the whole story
alvinashcraft
31 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

The Claude Science product guide

1 Share
The Claude Science product guide
Read the whole story
alvinashcraft
38 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Microsoft Research and Paige Unveil PRISM 2 Pathology AI Model

1 Share

Microsoft Research and Paige’s PRISM 2 combines nearly 2.5 million pathology images with clinical language, creating a versatile AI model for cancer detection, biomarker prediction, and research.

The post Microsoft Research and Paige Unveil PRISM 2 Pathology AI Model appeared first on Cloud Wars.

Read the whole story
alvinashcraft
3 hours ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories