Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
161882 stories
·
33 followers

Claude Sonnet 5.5

1 Share

Claude Sonnet 5.5

New Sonnet model from Anthropic today. They say it "runs 30%+ faster, and costs up to 30% less for most work" - it's priced the same as Sonnet 5 but appears to beat it on every benchmark, and should be cheaper to run as well.

Here are some pelicans riding bicycles. Sonnet 5.5 suffered from the same bug as Opus 5.5: the "max" thinking effort pelican thought for 128,000 tokens (at a cost of $1.28) before running out of tokens and failing to produce an SVG.

Here's the pelican it gave me for thinking effort "xhigh", at a cost of 5.74 cents and taking 41 seconds:

It's good- correct bicycle frame, legs either side of the frame, feet touching the pedals, chain in the right place, it is wearing a misshapen blue bicycle helmet though.

Sonnet 5.5 appears to be almost as good as Opus 5.5 on some coding tasks, including various viral 3D animation tricks.

The most interesting thing about Sonnet 5.5 is that it's now the model used for the free tier on claude.ai. OpenAI's ChatGPT free tier uses Luna 5.6, which means Anthropic currently have a much more capable free offering.

I ran this prompt against that free tier:

build me an HTML page that renders a three-dimensional pelican riding a bicycle using WebGL

And got back this page, which is a solid effort.

Anthropic's announcement reiterates that Haiku 5.5 will be available "in the coming weeks". I really hope that one is price-competitive with GPT-6 Luna!

Tags: ai, generative-ai, llms, anthropic, claude, pelican-riding-a-bicycle, llm-release

Read the whole story
alvinashcraft
5 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

When can we say AI made a scientific discovery?

1 Share

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here.

Last Wednesday, Anthropic announced that earlier this year it had launched a molecular biology lab, where Claude agents read and conjecture about hard biology problems and human scientists run experiments on what they report. And this AI-powered lab, the company said, had made its first discovery. 

To understand what Anthropic says its system did, imagine you’re flipping through a library of millions of DNA sequences, amassed as scientists sequence more and more of the living world. One step toward a breakthrough might be finding a peculiar sequence that encodes an interesting enzyme, perhaps. Then you’d need to figure out what that enzyme does and, eventually, how to manipulate it to do something useful.

What Anthropic says its system of 950 agents found after 21 hours was not a brand-new sequence. The agents instead flagged a repeating pattern surrounding a known enzyme, a particular pattern Anthropic said hadn’t been catalogued before. But if you read through Anthropic’s announcement, which calls this pattern “reminiscent” of what led to the gene-editing technology CRISPR that “has already transformed science and medicine,” it sounds as if this army of agents really found something of note. 

These claims have angered some biologists. A viral post from one, subsequently endorsed by the chair and CEO of the drugmaker Eli Lilly, said that “finding a weird cluster of genes and repeats is often the easy part. The hard part, and where the real discoveries come from, is figuring out what the system actually does.” The agents helped with some laboratory grunt work, in other words. But a discovery it is not. 

It’s a reminder that even if AI does something impressive—like finding a pattern in a mass of biological data that would be difficult to perceive with human eyes alone—the result itself may not constitute a breakthrough for science. What is novel for AI may be routine, unsurprising, or simply not that consequential to a biologist.

Muddying the issue further, Mario Rodríguez Mestre, a biologist at the University of Copenhagen, said over the weekend that his team had already discovered this particular pattern, the New York Times reported. Mestre, who regularly chatted with Claude in his work, wondered whether Anthropic’s team had learned from his conversations. Anthropic denies this, but Mestre says he’s stopping all use of Claude anyway.

Part of the problem here is that AI companies aren’t presenting their systems simply as tools scientists can use, like microscopes or supercomputers. They’re insisting that the AI systems are making discoveries themselves. To some, that approach is  incompatible with how science actually works, with new knowledge more typically emerging from collaboration and an ever-growing arsenal of tools. 

It’s also making people more skeptical of genuine progress when it happens. Whittling 200,000 candidates down to a few worth exploring is no small feat; it is legitimate scientific work. The fact that a general-purpose chatbot could do that work is notable, even if humans helped steer it and ultimately ran the experiments. But once the standard is whether Claude itself made a discovery, all that becomes evidence for one side or the other in a debate that has only two answers: breakthrough or bust.

Once we’re judging AI by whether it has made a discovery, it’s also tempting to shift the goalposts even after it really does seem to notch a win. Earlier this month, OpenAI said its own team agents had cracked a million-dollar problem in mathematics. But a couple of weeks later, nearly every AI skeptic in my feed was sharing an article asking whether it was the math problem that really mattered. 

To be clear, the piece did not argue that OpenAI’s solution was wrong. Instead, it argued that the particular result may not be the one mathematicians care most about. Throw in the accusation by a mathematician that the models may have used some of his work without credit, and people are left thinking either OpenAI cheated or the solution wasn’t important anyway. Or both.

That’s part of what concerns Lucas Harrington, the biologist who wrote the post critiquing Anthropic’s announcement. He closed with a suggestion: AI companies, he said, should “set the bar high now, so that when an AI actually discovers a fundamentally new biological mechanism, everyone appreciates how big a deal it is.” But as OpenAI’s Sam Altman and Anthropic’s Dario Amodei race to one-up each other, raising the bar for scientific breakthroughs by AI might be the last thing on their minds.

Read the whole story
alvinashcraft
5 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Are you a Codex Original?

1 Share
We’re collecting real stories of builders, tinkerers, researchers, and creators who are using Codex to do incredible things. If you want to be a part of the next chapter of the Codex Originals program, tell us more about your story and project below.
Read the whole story
alvinashcraft
6 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Giving companies more control over their AI agents, with NVIDIA

1 Share
Giving companies more control over their AI agents, with NVIDIA
Read the whole story
alvinashcraft
6 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Bill Gates Says an AI 'Kill Switch' Isn't Enough

1 Share
Bill Gates said an AI "kill switch" alone would be insufficient to address the technology's risks. "The kill switch is, that's kind of a weird thing. It's not enough to have a kill switch," the Microsoft co-founder said in an interview that aired Sunday on NBC's "Meet the Press." He said the more immediate concern is people using AI for cyberattacks or bioterrorism rather than systems acting autonomously. "We're not yet at the point where they autonomously grab computers and, you know, can't be shut down," he said, adding that the public discussion "just sort of shows how nontechnical various people are." Politico reports: Gates said he was not opposed to a kill switch but argued that companies should monitor sophisticated AI models and keep records of "exactly what's being done" to help guard against people using the systems for cyberattacks or bioterrorism. "The United States ought to make sure this is done. China will want to do this," he continued. [...] Gates also called for legislation requiring safeguards and monitoring, saying voluntary measures by AI companies were insufficient. "No one thinks self-regulation is enough," he said. He argued that the more immediate danger was people using AI to cause harm rather than AI systems acting on their own. "The thing that's urgent has to do with bad people using AI, not the AI going off on its own," he added.

Read more of this story at Slashdot.

Read the whole story
alvinashcraft
6 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Volkswagen replaces ID.4 with all-electric Tiguan

1 Share
VW ID.Tiguan

In a widely expected move, Volkswagen announced Monday that it will replace the recently retired ID.4 crossover with the upcoming ID.Tiguan. The decision is an acknowledgment by the German automaker that its more recognizable nameplates like Tiguan are likely an easier sell with consumers, especially when it comes to EVs.

When it lands in Europe in early 2027, the ID. Tiguan will join the ID.Polo and ID.Cross to round out Volkswagen's new-and-improved electric lineup. VW also confirmed that the ID.Tiguan will come to the US, where the ID.4 was previously manufactured, "at a later date." Design details remain under wraps, as the ID.Tiguan i …

Read the full story at The Verge.

Read the whole story
alvinashcraft
6 hours ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories