Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162016 stories
·
33 followers

Simplify Directory File Operations in .NET with Spargine’s DirectoryInfoExtensions

1 Share
The HttpRequestExtensions class in DotNetTips.Spargine.Extensions offers reusable methods to streamline HTTP request handling, enhancing code maintainability and consistency. Key functions include reading request bodies, managing headers, extracting bearer tokens, and validating content types. This utility supports input validation and unit testing, improving code quality in applications.
Read the whole story
alvinashcraft
12 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Deterministic When Possible. Probabilistic When Necessary. Human When Cheaper.

1 Share

One inescapable conclusion I’ve had to draw observing our industry over the last couple of years is that the AI thing isn’t about productivity. If it was, teams would be looking for ways to improve outcomes. But that’s not what I’ve been seeing.

I’ve seen dev teams setting out with the specific goal to use AI coding agents as much as possible, with the impact a secondary (or even a thirdary) concern.

I’ve watched developers prompting Claude Code to rename classes or methods when there’s a shortcut in their editor that will do it with a fraction of the keystrokes, using a tiny, tiny fraction of the compute and the energy, and do it more reliably.

I’ve watched them task Copilot with finding unused code when background compilation has already identified it.

I’ve watched them explain the code they want to Codex in more words than the code itself, and the model still gets it wrong.

Arguably these people have lost the plot.

And there are those other teams; the ones who use the best tool for the job. Need a refactoring? IntelliJ or Rider or PyCharm’s got you covered most of the time. Wanna know where the unused code is? Your IDE’s showing you. And if you don’t have that feature, a linter will do it lickety-split without the need for a £20,000 GPU and a terabyte of VRAM. If you know what code you need, maybe just write it. M’kay?

And then they hit a gap in their tooling. They need to move an instance method in Python. PyCharm doesn’t have that refactoring. So they go the agent window:

> Move the method calculateDiscount from the Order class to the Product class

And – 90% of the time – the model will do what they need. (And, annoyingly, sometimes more than they need.)

To perform the refactoring by hand would usually take longer, so they make a rational choice to throw the dice if it will save some time.

These are the teams whose outcomes have improved. Lead times and release cycles have shrunk a little, and it hasn’t been at the price of reliability.

And that’s what this is supposed to be about, right? Better outcomes. Bang for the buck. Or should I say “Bang for the token”?





Read the whole story
alvinashcraft
16 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

SvelteKit 3 Reaches Release Candidate, Moving Config to Vite and Retiring the $lib Alias

1 Share

The Svelte team has entered the release candidate phase for SvelteKit 3, which focuses on code refinement and prepares for future updates. Notable changes include moving configuration to vite.config.ts, replacing the $lib alias with #lib for improved compatibility. SvelteKit 3 requires Vite 8 and Svelte 5, enhancing error handling and build efficiency, while retaining some experimental features.

By Daniel Curtis
Read the whole story
alvinashcraft
34 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Upgrading Dependency Track from v4 to v5

1 Share

Dependency Track is an open source component analysis platform from OWASP. You upload the SBOM of your application, and Dependency Track keeps track of the components inside it. It checks those components against vulnerability sources like the National Vulnerability Database, GitHub Advisories and OSV. It also lets you define policies and send notifications when something new shows up. In short: it tells you which of your applications are affected when the next vulnerable library hits the news.

Recently Dependency Track got an upgrade and version 5 was released. So, time to upgrade!

However, that turned out not to be as easy as expected. It took us 2 attempts. Our first attempt failed completely, so we took a different route. Here is what we tried and the two things that cost us the most time.

Big shout out to Jef, who looked over my shoulder during the upgrade and helped tackling the issues when we got stuck.

Two approaches in the documentation

The Dependency Track documentation describes two ways to get a new version running.

  • Upgrading running instances: the in-place approach. You run the schema migrations, then replace the API server instances one at a time. This gives you no planned downtime, as long as more than one instance is running and the release notes don't call for a full stop.
  • Migrating from v4 to v5: a separate migration. The v4-migrator tool copies your v4 data into a new v5 database. v4 has to be offline while it runs, and v5 only supports PostgreSQL.

We started with the first one. That turned out to be no success for us.

Attempt 1: the in-place upgrade

We pointed the v5 container at our existing v4 database and let it migrate on startup.

That didn't work. The migration process started, but failed with the following error message:

Caused by: org.flywaydb.core.internal.sqlscript.FlywaySqlScriptException: Failed to execute script V202605111028__add_latest_version_published_at_column_to_package_metadata.sql

No matter what we tried, we couldn't fix this error.

Remark: v5 renames many configuration properties and environment variables, and it no longer accepts the v4 names. If you reuse your v4 configuration in Azure Container Apps, check it against the v5 configuration reference.


Attempt 2: migrate the database separately

Instead of one big step, we split the upgrade in two. First, we migrated the data into a v5 database with v4-migrator, then we started the v5 application on top of it.

Before you start, the documentation lists a few requirements:

  • v4 must be on version 4.14.2 or later, and the v4 API server must be stopped.
  • The target is a dedicated PostgreSQL 14+ database. Don't start the v5 API server against it before the migration has run.
  • The database user you use for the migration should be the same one the v5 API server uses afterwards.

The migration runs in steps: bootstrap applies the v5 schema on the target database, verify checks the target, and run does the extract, transform and load. Afterwards you verify again and cleanup the staging schema.

The migration tool is available in a docker container that we ran locally during the upgrade process.

Two things cost us a lot of time.

Caveat 1: the quotes in the documentation

The documentation shows the JDBC URLs in single quotes:

docker run --rm -t ghcr.io/dependencytrack/v4-migrator:5.0.0 bootstrap \
  --target-url 'jdbc:postgresql://target-host:5432/dtrack' \
  --target-user dtrack \
  --target-pass

That doesn't work for us. It only worked after we removed the quotes:

docker run --rm -t ghcr.io/dependencytrack/v4-migrator:5.0.0 bootstrap \
  --target-url jdbc:postgresql://target-host:5432/dtrack \
  --target-user dtrack \
  --target-pass

The error didn't point at the quotes, so we checked permissions, connectivity, versions (and even the weather) first.

Caveat 2: the pg_trgm extension

Dependency Track v5 needs the pg_trgm extension. The bootstrap step tries to install it, and the preflight check verifies it.

This extension was not installed out-of-the-box. So, the preflight check failed. Because we were using an Azure Database for PostgreSQL Flexible Server: the extension must be allow-listed first through the azure.extensions server parameter:

After that, we could create the extension:

CREATE EXTENSION IF NOT EXISTS pg_trgm;

Starting the v5 application

With the migrated database in place, we could finally deploy the v5 container to Azure Container Apps.

This time it started without problems.

Tip: The migration is deliberately lossy in a few places. Every notification rule is disabled after the migration, and repository and analyzer credentials have to be entered again in the v5 secret manager. Plan some time for this after the cutover.

That's it!

More information

Read the whole story
alvinashcraft
46 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Anthropic says it fixed Claude's writing. I ran the evals to check.

1 Share

TL;DR: Anthropic says Opus 5.5 fixed Claude’s writing. I tested that claim with the standard eval loop in Arize AX: build a dataset, run an experiment for each model, annotate the output by hand, then build an evaluator from the annotations and run it.

The em dash really is gone: 12.9 per 1,000 words in Opus 5, and two in the entire 57,000 words of Opus 5.5 output. The rest of the Claudisms halved. Better, but not fixed.

I’m sick of Claudisms. You know the ones. You’re scrolling LinkedIn or X, and a post tells you that something “isn’t just a tool, it’s a mindset,” or that a detail “is doing a lot of work,” and that you should sit with that. Humans don’t really write like that, but Claude does, and so does everyone who pastes Claude’s output straight into a post.

When Opus 5.5 came out, one of the launch details caught my eye. The Opus 5.5 announcement says, “We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5,” and that it “is less likely to use jargon or idiosyncratic phrases.” Anthropic staff went a step further. Sholto Douglas posted, “We fixed the writing,” and Tom Brown posted, “we fixed the accent.” The Decoder’s headline said Anthropic “promises less ‘Claudish’ writing.”

That’s a testable claim. And I work at Arize, so when someone says a model got better at something, my first reaction is “show me the eval.” Vibes don’t count, whether they’re mine or a vendor’s.

So I built claude-compare, and ran the same loop I’d use to test any change to an agent: build a dataset, run experiments on it, annotate the output, then build an evaluator and run it.

The eval loop: build a dataset, run experiments, annotate the output, build an evaluator, run it, then iterate on the evaluator against the annotations

Step 1: Build a dataset

A dataset is the fixed set of inputs you test against. Every experiment runs over the same examples, so when a score moves, you know it was the thing you changed and not the inputs.

For a writing test, the inputs are writing tasks. I used 20 research briefs, each one containing the notes for a blog post, so that I had long form content to review for Claudisms. The topics cover five genres (opinion, technical explainer, tutorial intro, news analysis, product announcement) and five domains (AI, travel, cooking, games, books). A dataset of AI topics alone would only tell you how Claude writes about AI.

Two things make a dataset like this trustworthy:

It’s frozen. Opus 5 researched each topic once, with web search, and the briefs are committed and hash-locked. The writer gets no tools and no network, so the brief is all it has to go on.
The inputs don’t carry the style you’re testing for. The briefs are written as bullet fragments, so Opus 5’s style can’t leak into the input.

The brief is the load-bearing part of the whole setup (sorry).

Step 2: Run an experiment for each model

An experiment is one run of your task over the whole dataset, with each output stored against the example that produced it. When you compare two experiments, you’re comparing whatever differs between them, so you want that to be one thing.

The claude-compare-full-v1 dataset in Arize AX with its four experiments

If you want to compare how two models write, the first step is to just change the model. Sounds obvious, but folks often mix the model change with a prompt change as well. Changing the model first, measuring, and then changing the prompt if necessary, gives you a much clearer view on the impact of your changes.

That means pinning everything else. Both models get an identical prompt. Effort is pinned to medium, because Opus 5 defaults to high and Opus 5.5 to medium, and leaving the defaults in place would have measured two different settings. The harness also checks the served model on every response, so each experiment really is the model it says it is.

Model output varies from run to run, so I ran each model over the dataset twice. That gave four experiments (two models, two repeats), 80 posts and about 110,000 words. In AX the four experiments sit on the same dataset, so every evaluator I add later scores them all the same way and I can compare Opus 5 and Opus 5.5 side by side.

Step 3: Annotate the output

Annotations are human labels on experiment output. They’re your ground truth: what a perfect evaluator would say. You want them before you build the evaluator, because they tell you what you’re actually measuring and they give you something to check the evaluator against later.

You should label at the level you want the evaluator to work at. A score per post would tell me which posts felt Claude-ish, but not why. Marking individual phrases tells me which sentences, and an evaluator that returns phrases can be checked line by line.

The labelling tool doesn’t need to be fancy. I put all 40 Opus 5 posts into one Google Doc and read the lot, all 53,000 words, leaving a “Claudism” comment on every phrase that made me wince. That came to 153 flags across 38 of the 40 posts. The flags then went into AX as annotations on the Opus 5 runs, so they sit next to the output they describe.

Jim’s “Claudism” comments down the margin of the Opus 5 review doc

Then I looked at what I’d flagged, and it wasn’t what I expected.

Jim’s 153 flags grouped by move, with the nine the stock-phrase list caught

I had a hand-written list of the famous Claudisms: “load-bearing”, “delve”, “crucially”, “genuinely,” and so on. It matched just nine of my 153 flags. “Load-bearing,” the phrase everyone jokes about, appears three times in 53,000 words of Opus 5.

What I’d actually flagged were rhetorical moves. The biggest group, about 48 of the 153, was the text telling me something was important instead of showing me why: “The interaction matters,” “a fact worth internalising,” “deserves a moment.” Next were contrast reframes (“It’s a topology, not a genre.”), verdict intensifiers (“the honest answer,” “the whole point”), and signposts that tease an insight instead of giving it (“Here’s the part that surprises people…”).

None of those are fixed wording, so no phrase list will ever catch them reliably. It’s not a phrase problem. It’s a moves problem. 🫠

Step 4: Build an evaluator from the annotations

An evaluator scores every run in an experiment automatically, so you don’t have to read 60,000 words again each time something changes. The annotations shape how you build the evaluator.

The claudism_spans evaluator in Arize AX

Start with the output format. My first evaluator was an LLM judge that gave each post a single 1 to 5 score for how Claude-ish it read. That’s easy to build and almost impossible to check. If it says a post is a 4, which sentences made it a 4? You can’t line a single number up against 153 human flags. So the version I kept is a span judge: it returns every Claudism it finds as an exact quote, in the same shape as my annotations, and I can match each quote against my flags to measure recall directly.

The categories come from the annotations too. The judge sorts each quote into one of six moves taken from my flags: salience flag, contrast reframe, verdict intensifier, signpost, gotcha framing (“the trap is”) and stock metaphor (“load-bearing”, “earns its keep”). And the annotations double as test cases. All three “load-bearing” sentences have to be caught, or the run fails.

A few other choices apply to almost any evaluator:

  • Use code when you can. Counting em dashes doesn’t need an LLM, so that’s a plain code evaluator with no wiggle room.
  • Don’t let a model grade its own family. The judge is OpenAI’s gpt-6-luna, so Claude isn’t marking Claude’s homework.
  • Normalise for length. Opus 5.5 writes about 8% longer, so the judge’s quotes become Claudisms per 1,000 words rather than raw counts.

Both evaluators run in AX against all four experiments.

Step 5: Run it

The em dash really is gone. Opus 5 uses 12.9 em dashes per 1,000 words. Opus 5.5 used two in its entire 57,000 words of output. As Brodie Robertson put it, “The em dashes have been deleted I repeat the em dashes have been deleted.” That result comes from a code evaluator, so there’s no judge involved and no wiggle room.

Em dashes fell from 12.9 to 0.05 per 1,000 words, and Claudisms fell from 4.79 to 2.38

The other Claudisms halved. The span judge found 4.79 Claudisms per 1,000 words in Opus 5 and 2.38 in Opus 5.5, a 50% drop, and every category went down:

Category Example Opus 5 Opus 5.5 Change
Salience flag “This matters.” 1.67 0.93 −44%
Verdict intensifier “The honest answer is…” 1.10 0.38 −65%
Signpost “Here’s the part that…” 0.79 0.53 −33%
Contrast reframe “It’s a topology, not a genre.” 0.63 0.27 −56%
Stock metaphor “load-bearing”, “earns its keep” 0.43 0.13 −70%
Gotcha framing “The trap is…” 0.16 0.13 −19%
All Claudisms 4.79 2.38 −50%

Claudisms per 1,000 words, from the span judge running in Arize AX.

Salience flags, the “this matters” move, are still the most common Claudism in both models. Gotcha framing is too rare to read much into, with 8 uses against 7. The gap holds in every genre and every domain I tested.

Opus 5 and Opus 5.5 experiments compared side by side in Arize AX, with the em dash and Claudism evaluations

Opus 5.5 also picked up some new habits. A simple scan found it uses about 2.5 times as many bold lead-in bullets as Opus 5, 35% more three-part lists, and writes 8% longer. So some of the old tics were traded in rather than dropped.

My favourite detail: Anthropic’s own prompting guide for Claude Fable 5.1 warns about “mannered prose,” using “this point earns its keep” as its example. Opus 5.5 still wrote that a stand mixer “earns its counter space” and that a searing technique “earns its place.”

So the claim holds up halfway. “Fixed” is too strong. “Much better” is fair.

Iterate on the judge prompt

The 50% figure came from the third version of the judge, not the first. An evaluator is a prompt like any other, and you rarely keep the first draft of a prompt. The annotations from step 3 are what let me iterate on it with numbers instead of gut feel.

Three versions of the judge on the same 80 posts: a 27%, 34% and 50% drop, with the share of real Claudisms in a spot check under each

The first version said Claudisms fell by only 27%. Moving the judge to gpt-6-luna gave 34%. Recall against my flags was high on both runs, so the judge was finding the Claudisms I’d marked.

Here’s the part that bites people. (Yes, I know, sorry.) Recall only tells you what the judge caught. It says nothing about what else it tagged. So I read a random sample of the spans it found in the Opus 5.5 posts. Plenty of them were ordinary writing. “First, a definition.” got tagged as a signpost. “Strain the context window” got tagged as a stock metaphor. “It applies to all output tokens, not only thinking.” got tagged as a contrast reframe, when it’s just being precise.

Only about 40% of the Opus 5.5 spans in that sample were real Claudisms. (That spot check was done by an AI and is small, so take it as rough). The false positives weren’t random noise. Every writer uses plain transitions and ordinary metaphors, human or model, so they formed a floor under both scores and made the two models look closer than they are.

Sit with that for a minute. (Sorry.)

The third version gave the judge a test to run on every phrase before it tags it: imagine deleting the phrase and reread the sentence. If the post loses information, such as a fact, a number, how something works or what to do, the phrase carries information and it isn’t a Claudism. If nothing is lost, it’s filler, and it counts. Delete “This matters.” and the post says exactly the same thing, so it gets tagged. Delete “not only thinking” from “It applies to all output tokens, not only thinking.” and you lose the point of the sentence, so it doesn’t.

I also added examples of what not to tag in each category and limited stock metaphors to the well-worn ones. Precision on the Opus 5.5 sample went up to about 26 in 30. Recall against my flags on held-out posts dropped from 74% to 64%, and it still caught all three “load-bearing” sentences. I’ll take that trade: an evaluator that misses a few real Claudisms beats one that counts plain English as a Claudism.

Each version ran against the same annotations, so every change came with a number for what it gained and what it cost. I tuned on half the posts and held the other half back to check the result. Without the annotations, the 27% would have looked like a perfectly good answer.

What I’d tell you before you try this

If you’re writing with Claude, hunt your own drafts for the Claudisms, more than the em dash or “delve”: telling the reader something matters, knocking down a claim nobody made, teasing a point instead of making it.

The code, the categories and the judge prompt are all in the claude-compare repo, along with a de-styling tool that uses the same categories as an editing brief.

The bigger lesson applies well beyond Claude. When someone tells you a model got better, don’t trust the vibe, and don’t trust a vendor’s word for it. Build a dataset, run the experiment, annotate the output, build an evaluator from the annotations, and iterate on it against those annotations until you trust its numbers.

And this matters. 🙃

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete

As a general rule, calling product support while drunk is not recommended

1 Share

In a reminiscence about product support stories, a now-retired colleague related a story of a call that they took when working the product support phone lines for Microsoft FrontPage, a Web site authoring tool from the late 1990’s and early 2000’s.

The call came from a person who was creating a fan site for the rock band Van Halen. They had a java applet on the web site that played music, this being back in the days when java applets on Web pages was a thing, and the applet wasn’t working.

This was not a FrontPage issue, but my colleague figured they could try to help out anyway. “But man, was this person lit.” The sound of ice clinking in a glass was readily apparent, and as the call went on, the customer got more and more drunk.

The customer shouted in frustration, “You’re not even a real f—ing fan of Van Halen!”

My colleague advised the customer that if the swearing continued, they would have to end the call.

This was apparently not the response the customer was hoping for.

“F— you! I wanna talk to Gates!”

I don’t know what happened next. I wonder if the customer was transferred to the special operators who pretend to be Bill Gates’ secretary.

The post As a general rule, calling product support while drunk is not recommended appeared first on The Old New Thing.

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories