Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162513 stories
·
33 followers

WW 1004: Fat-Washed - RTX Spark and Hybrid AI

1 Share

Microsoft and NVIDIA's RTX Spark laptops are official, with preorders open and shipping October 16. Prices run from $2,599 (24GB) to $5,899 (128GB, already sold out on Surface). Paul and Leo cover who needs that much RAM, what "hybrid intelligence" means, and why Paul's new Googlebook disappointed him.

Windows

  • RTX Spark laptops: from $2,599 (24GB), $3,700 (32GB), $5,899 (128GB). Ships Oct 16
  • Partners: Microsoft, ASUS, Dell, HP, Lenovo, MSI. Paul reviews the HP OmniBook Ultra 16 first
  • Surface RTX Spark Dev Box: $6,000+
  • "Hybrid intelligence": local AI first, cloud when needed
  • Windows Search gets inline AI actions in Insider builds
  • Cloud Rebuild adds better remove and sanitize options
  • Windows app adds Remote PC connections (preview)
  • Intel Wildcat Lake: Paul's test laptop was slow even with 16GB
  • Googlebook first look: Lenovo Googlebook 15 (~$1,100) is missing basics an iPad has
  • Microsoft has tried to kill Windows many times, including the leaked Project Aion

AI/Dev

  • Microsoft AI ships new voice and transcription models
  • Gemini 4 Argon is announced, but limited to cybersecurity partners
  • Google Docs and Drive add native Markdown
  • C# Dev Kit 11 arrives with .NET 11

Xbox and Gaming

  • Rockstar corrects reports of exclusive Xbox GTA VI streaming. Launch is Nov 19
  • Xbox Elite Series 3 controller leaks again, with a built-in screen
  • October Game Pass: Battlefield 6 and more
  • Sony brings QSSR AI upscaling to the standard PS5
  • Google Playground lets you vibe-code games

Tips and Picks

  • Tip of the week: Desktop98.com, a Windows 98 simulator in your browser
  • App pick: Gears of War: E-Day (Game Pass Ultimate and PC)
  • Brown liquor pick: 1792 Small Batch bourbon

Hosts: Leo Laporte and Paul Thurrott

Download or subscribe to Windows Weekly at https://twit.tv/shows/windows-weekly

Check out Paul's blog at thurrott.com

The Windows Weekly theme music is courtesy of Carl Franklin.

Join Club TWiT for Ad-Free Podcasts!
Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit

Sponsors:





Download audio: https://pdst.fm/e/pscrb.fm/rss/p/mgln.ai/e/294/cdn.twit.tv/cap/ww_1004/bcbbae1d-81c4-43ea-b115-7eb58f6341d6.mp3
Read the whole story
alvinashcraft
15 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Cloudflare Open Sources Decision Models for AI Agents

1 Share

During its recent “Birthday Week”, Cloudflare announced Clef, a set of open-weight AI models designed to choose between predefined options rather than generate text. Cloudflare released 9B- and 27B-parameter models, along with a platform for adapting them to specific decision-making tasks.

By Renato Losio
Read the whole story
alvinashcraft
36 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Requirements Traceability Matrix for AI Software Delivery—From Issue to Pull Request

1 Share

Your issue tracker and version control platform hold most of the requirements traceability matrix an auditor wants, in a form nobody in your org can query.

On a Tuesday, an auditor asks a simple question. Which requirement produced the code in release 4.2? Who verified that it worked? Instead of answering from Jira and GitHub, someone spends two days reconstructing the story from merged pull requests and old release notes. It’ll still be a best guess, delivered with total confidence.

You need a requirements traceability matrix. It connects each requirement to the work that implements it, then traces every delivered artifact back to the reason it exists. Its purpose is to make the auditor’s question boring.

In most engineering orgs, it does the opposite, because someone has to keep it current by hand. That manual work doesn’t scale, and nobody was ever promoted for updating it. Meanwhile, your development workflow is already creating the links the matrix needs, inside systems you already pay for, where nobody can query them.

What the Matrix Was Supposed to Do


Image generated with AI

The job of the matrix is to connect the reason for a change to the issue and pull request that carried it. It also ties that work to the test run and approval that checked it. It answers two questions. Given a requirement, what proves it shipped? Given a diff, what business reason justifies it?

That first question matters outside an audit. A team that can trace a requirement to the work behind it can understand unfamiliar code faster.

But the matrix usually gets treated as audit paperwork. When that happens, it never earns budget. It runs on whatever time is left over, and that time keeps failing to materialize.

The research says underfunding traceability costs more than it saves. A 2015 controlled experiment gave participants real maintenance tasks on unfamiliar projects. Some tasks included traceability links between requirements and code. Participants who could follow those links finished 24% faster on average and produced 50% more correct solutions.

Faster work frees capacity, and capacity determines what your team can take on. Traceability, it turns out, is a productivity feature wearing a compliance costume.

The Chain Your Workflow Leaves Behind

Set the hand-kept matrix aside and trace a change through your org. Each handoff leaves a record whether anybody wants it to or not:

  • Requirements arrive from customer commitments or regulations.
  • A Jira or Azure DevOps issue names the requirement.
  • The implementation plan lives in the issue description or a linked design doc.
  • Commits on a branch reference the issue by number.
  • CI runs unit tests and acceptance tests against those commits.
  • A reviewer approves or rejects the pull request, on the record, with a timestamp.
  • The pull request merges and, closing keyword satisfied, closes the issue it named.

Requirement to issue to plan to code to tests to review to merge.

LinkUsual sourceProof
RequirementProduct brief or contractWhy the work exists
IssueJira or Azure DevOpsWhat work was approved
PlanIssue description or design docHow the change should be made
Code changeCommit or pull requestWhat changed
Test resultCI log or test reportWhether verification passed
ReviewPull request approvalWho accepted the change

How much of that chain your platforms hold varies by team. One team writes fixes #4127 in every commit. The next puts PROJ-4127 in the pull request title. A third uses only a branch name like adam/auth-timeout-fix, because its tech lead is confident the branch name speaks for itself. It doesn’t.

Those references don’t join unless your teams use the same pattern. GitHub accepts several closing keywords, for example. Picking the one every team uses and enforcing it with a required status check is this quarter’s decision. It’s a boring decision, which is why it keeps not getting made.

Agent Volume Breaks the Hand-Kept Matrix

Now put coding agents into that workflow. The number of branches and pull requests rises, while human memory of why each change happened stays thin. The hand-kept matrix stops getting updated because somebody has to open it and add the link. That somebody has a release to ship, and the matrix loses that argument every single time.

Agent-authored change puts three questions on your desk. Who reviewed the work? Where is the agent’s implementation plan stored? What Git author or pull request author identifies the agent and the human who directed it? A shared bot account answers none of this, though it does look tidy in the commit log.

The honest objection is that this chain proves custody, not correctness. Commit text like fixes #4127 proves only that somebody typed a number. Tests and review still carry the correctness question. GitLab, for instance, removes existing approvals when new commits land on the source branch by default, so an agent push resets them.


Traceability Check: Pick one merged pull request. Working from Jira and GitHub alone, name the requirement it served. Then name the person who approved it. If you have to ask a developer, that record does not exist.


Your CI Retention Expires the Evidence on a Timer

Agent volume is not the only clock running on the evidence chain from issue to merge. If you ship a high-risk AI system into the EU market, Annex IV of the EU AI Act requires technical documentation carrying test logs and reports dated and signed by the responsible persons, and Article 11 requires it stay current.

Staying current under a rule like that assumes the evidence still exists. Test runs and approvals are worth only what your platforms keep. Most hosted CI platforms delete workflow logs and test reports after a default retention period nobody in your org chose.

GitHub is the worked example. It retains workflow artifacts and logs for 90 days by default, adjustable up to 400 days on private repositories. The JUnit report and workflow log for a March release can be gone by June unless somebody changed a setting that has never once come up in a planning meeting.

Months later, your biggest customer’s security team asks for the verification record of one release. The issue and pull request are still there. The test run was deleted 90 days after it passed, and the reviewer who approved the code has left. You’re attesting to work you can no longer show, which is a memorable position to negotiate a renewal from.

Traceability Is a Read Problem

Retention keeps the evidence alive; finding it is a different problem. Most traceability programs fail on the write side, which your workflow already handles. What your org lacks is a read path.

One query should answer: show me requirement REQ-128, the issue that implemented it and the pull request that merged it. Then show the tests that passed and the reviewer who approved it.

Software supply chains already treat machine-emitted records as proof. SLSA provenance makes the build platform record how an artifact was made. Human-written claims about a build stay claims; machine-emitted ones carry their own proof.

Your workflow leaves the same kind of records behind. Building the read path means a scheduled job that calls the Jira and GitHub APIs, then stores issue IDs and approval timestamps.

Raise your CI artifact retention window this week, before you scope that job. On a private repository, that one setting buys 310 more days of evidence, which is the best return on a mouse click you’ll get this quarter.


Explore Progress Forge Orchestration

Turn AI-assisted development into a repeatable engineering process by orchestrating the coding agents you already use through structured, configurable development workflows—from work item to pull request.

Your coding agents execute the work. You own the process. Forge orchestrates it.

Try Forge Now

 

Read the whole story
alvinashcraft
42 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Getting started with DSC 3.0 – Part 5: Guarding your configuration with Assertion

1 Share

At the end of the previous post I said that dependsOn only decides the order in which resources are processed. It doesn't check that the resource you depend on is actually in the desired state.

Sometimes that's not enough. You want to say "only touch this machine if it's Windows Server 2025", or "only configure the website if IIS is really installed". That's what the Microsoft.DSC/Assertion resource is for.

Note: this post is part of a bigger series on Microsoft Desired State Configuration.

What is an assertion?

An assertion is a group resource. It contains a nested configuration document and runs the test operation on every resource inside it. It never changes anything.

  • If every nested resource is in the desired state, the assertion passes.
  • If one of them isn't, the assertion fails, and DSC doesn't invoke the resources that depend on it.

Only the resources that dependsOn the assertion are affected. The rest of the document runs as usual.

Remark: resources that only implement get, like Microsoft/OSInfo, are a natural fit for assertions. They can read the state but not change it.

A minimal example

Save this as assertion.dsc.yaml:

$schema: https://aka.ms/dsc/schemas/v3/bundled/config/document.json
resources:
- name: Demo key
  type: Microsoft.Windows/Registry
  properties:
    keyPath: HKCU\Software\Demo
    _exist: true
  dependsOn:
  - "[resourceId('Microsoft.DSC/Assertion', 'Windows only')]"
- name: Windows only
  type: Microsoft.DSC/Assertion
  properties:
    $schema: https://aka.ms/dsc/schemas/v3/bundled/config/document.json
    resources:
    - name: Operating system
      type: Microsoft/OSInfo
      properties:
        family: Windows

Two things to notice:

  • The assertion has its own $schema and resources, nested inside properties
  • The registry key depends on the assertion with the same resourceId() syntax we used in Part 4, with Microsoft.DSC/Assertion as the type and the instance name

Apply it:

dsc config set --file assertion.dsc.yaml

The operating system is Windows, so the assertion passes and DSC creates the key. 

Check it:

Test-Path HKCU:\Software\Demo


Let it fail

Remove the key first:

Remove-Item HKCU:\Software\Demo

Now change family: Windows to family: Linux in the document and run the same command again:

dsc config set --file assertion.dsc.yaml

The assertion fails, so DSC doesn't invoke the registry resource. Run Test-Path again and the key is not there.

Remark: the registry resource is listed before the assertion in the document. Like we saw in Part 4, the order in the document doesn't matter. The dependency does.

A more realistic case

A common use is applying settings depending on the Windows Server version. The OS version check is the gate:

$schema: https://aka.ms/dsc/schemas/v3/bundled/config/document.json
resources:
- name: Baseline tag
  type: Microsoft.Windows/Registry
  properties:
    keyPath: HKLM\SOFTWARE\Contoso
    valueName: Baseline
    valueData:
      String: Server2025
  dependsOn:
  - "[resourceId('Microsoft.DSC/Assertion', 'Server 2025 check')]"
- name: Server 2025 check
  type: Microsoft.DSC/Assertion
  properties:
    $schema: https://aka.ms/dsc/schemas/v3/bundled/config/document.json
    resources:
    - name: Server 2025
      type: Microsoft/OSInfo
      properties:
        version: "10.0.26100"

Add one pair like this per OS version and each machine only gets the settings that fit.

Here the version is an exact match. DSC 3.3 added version comparison to Microsoft/OSInfo, so you can assert a constraint instead. 

Remark: this one writes to HKLM, so run it from an elevated terminal.

Back to our previous example

In Part 4, dependsOn made sure IIS was installed before the web resources. But dsc config test on a clean machine could still fail, because the IIS cmdlets weren't there yet.

An assertion that checks the Web Server role fixes that. Add this to the document from Part 4:

- name: IIS is installed
  type: Microsoft.DSC/Assertion
  properties:
    $schema: https://aka.ms/dsc/schemas/v3/bundled/config/document.json
    resources:
    - name: Web server role
      type: PSDesiredStateConfiguration/WindowsFeature
      directives:
        requireAdapter: Microsoft.Adapter/WindowsPowerShell
      properties:
        Name: Web-Server
        Ensure: Present
  dependsOn:
  - "[resourceId('PSDesiredStateConfiguration/WindowsFeature', 'IIS')]"

And let the web resources depend on the assertion. For the application pool:

  dependsOn:
  - "[resourceId('Microsoft.DSC/Assertion', 'IIS is installed')]"

Give the website and the web application the same dependency, next to the ones they already have.

The flow is now: install IIS, check that IIS is installed, and only then configure the pool, the site and the application. If the check fails, DSC doesn't touch them.

Things to keep in mind

  • An assertion only gates the resources that depend on it
  • It never changes anything. If you want DSC to fix something, that's a normal resource
  • The assertion itself must be a top-level instance. A resource inside a group can only depend on its neighbors in the same group, so the IIS install and the assertion stay at the top

We are not there yet. In our next and probably last post about DSC, we'll add AI into the mix. 

Read the whole story
alvinashcraft
54 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

How can undefined opcodes ud0 and ud1 have parameters? How undefined were they?

1 Share

Gunnar Dalsnes wondered how ud0 and ud1 could have parameters if they were undefined? “Does it mean they were not completely undefined, just undocumented and not completely implemented?”

The opcodes ud0 and ud1 have no definition, but that’s not the same as being architecturally an “undefined instruction”. They live in a purgatory where they were not assigned a meaning, but were also not officially declared to be meaningless.

Originally, these byte sequences went through the instruction decoder and happened to slip through a few cracks before somebody finally noticed, “Wait a second, I don’t know how to execute this.”

You can see this when you look at the instructions that are encoded as 00001111 11111xxx, as of the Pentium III.

Bits Bytes Opcode Operand 1 Operand 2 Meaning
00001111 11111000 0F F8 PSUBB mm mm/m64 Subtract packed bytes
00001111 11111001 0F F9 PSUBW mm mm/m64 Subtract packed words
00001111 11111010 0F FA PSUBD mm mm/m64 Subtract packed dwords
00001111 11111011 0F FB No meaning assigned
00001111 11111100 0F FC PADDB mm mm/m64 Add packed bytes
00001111 11111101 0F FD PADDW mm mm/m64 Add packed words
00001111 11111110 0F FE PADDD mm mm/m64 Add packed dwords
00001111 11111111 0F FF No meaning assigned

The byte sequences 0F FB and 0F FF had yet to be assigned a meaning. You can see how the instruction decoder could take a shortcut and say, “Well, all the instructions in this range, or at least all the ones that I care about, take an mm registers and an mm/m64 operand, so I’ll just save myself some transistors and decode all of them with two parameters (mm, mm/m64).” And then after decoding, it would use bit 3 to decide whether to set up the arithmetic unit for an add or subtract, and it would use bits 0 and 1 to decide how to subdivide the bits into saturating units.

And if you gave it a 0F FF, it would be only in that last step that the decoder would realize “Oh dear, I don’t know what to do with a bit combination of 11. I’ll raise an invalid opcode instruction.”

The invalid opcode instruction got raised after the operands were parsed.

You can see the trouble that 0F FF created when those empty slots started to get filled in by the SSE instructions.

Bits Bytes Opcode Operand 1 Operand 2 Meaning
00001111 11111000 0F F8 PSUBB mm mm/m64 Subtract packed bytes
00001111 11111001 0F F9 PSUBW mm mm/m64 Subtract packed words
00001111 11111010 0F FA PSUBD mm mm/m64 Subtract packed dwords
00001111 11111011 0F FB PSUBQ mm mm/m64 Subtract packed qwords
00001111 11111100 0F FC PADDB mm mm/m64 Add packed bytes
00001111 11111101 0F FD PADDW mm mm/m64 Add packed words
00001111 11111110 0F FE PADDD mm mm/m64 Add packed dwords
00001111 11111111 0F FF I want to put PADDQ here but I can’t

the natural place to put the PADDQ instruction is 0F FF, but people had already been using 0F FF with the expectation that it raises an illegal instruction exception. Making it a valid instruction would break those programs, so Intel had to move PADDQ to the rather awkward location 0F D4.

Bonus chatter: Undefined instructions with parameters are actually not uncommon. For example, on AArch64, there is a range of 65,536 instructions set aside as permanently undefined, so the udf instruction takes a 16-bit immediate to specify which invalid opcode you want. The PDP-10 reserved opcode 000 as a permanently illegal instruction, and it carries a register and a memory address as parameters. (Because all PDP-10 instructions carry a register and a memory address as parameters.)

There are also so-called “unofficial opcodes” which are instructions that are not part of the instruction set architecture, but for which people reverse-engineered a consistent behavior and began to rely on it. (The 6502 processor is well-known in nerd circles for having undergone this type of analysis.) The 0F FF is one of these “unofficial opcodes” that was popular enough that Intel felt pressure to maintain backward compatibility with it, even though it was never architecturally documented or supported.

Bonus bonus chatter: It appears that the mnemonic ud1 was introduced by the nasm assembler:

* Added the following new instructions: SYSENTER, SYSEXIT, FXSAVE,
  FXRSTOR, UD1, UD2 (the latter two are two opcodes that Intel
  guarantee will never be used; one of them is documented as UD2 in
  Intel documentation, the other one just as "Undefined Opcode" --
  calling it UD1 seemed to make sense.)

It seems obvious that ud1 is also the name that Intel gave internally to that legacy instruction. Otherwise, there would be no need to call the new one ud2!

The post How can undefined opcodes <CODE>ud0</CODE> and <CODE>ud1</CODE> have parameters? How undefined were they? appeared first on The Old New Thing.

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete

Tokens and the Context Window: Out of the Window, Out of Mind

1 Share

Part two of ten. Most of the surprises people hit with AI at work trace back to this one chapter. A model reads in tokens, holds only so many at once, and charges for each one.

Figure showing that the model sees only what fits in its context window, and everything else is gone.

1. What is a token?

A token is the unit of text a model reads and writes. It can be a whole word, part of a word, a chunk of a number or a punctuation mark. For common English text, a rough rule is about four characters per token.

Why it matters. Tokens are how models are measured and billed. Context limits, prices and speed are all counted in tokens. Code and SQL use more tokens than you’d guess, because a name like SalesOrderDetailID splits into several pieces.

Watch out. The four-characters rule is a rough guide for English. Other languages, numbers and code can run much higher. When the number matters, use the vendor’s own token counter.

2. Why does AI struggle to count letters?

Because a token isn’t a word. A short, common word is usually one token, while a long or rare word splits into several. The model sees a few tokens, not ten letters, which is why it can miscount the letters in strawberry.

Why it matters. The same blind spot shows up in data work. Character counts, string lengths and exact spelling checks are weak spots. Let SQL or a script do the counting, and let the model explain the result.

The follow-up. “Why does it get simple arithmetic wrong sometimes?” Numbers are split into tokens too, and the model predicts digits rather than calculating them.

3. What is a context window?

The most text a model can consider at once, measured in tokens. It holds the instructions, the conversation so far, any files you attached and the answer being written. Anything outside the window doesn’t exist for the model.

What the interviewer wants. That on most models the answer counts against the same window. A long input can leave too little room for the reply.

Watch out. Window sizes change with every model release, so don’t memorise a number. Say how you’d find it: the model’s documentation lists it.

4. Why does AI lose track of a long file?

Fitting in the window doesn’t mean every detail gets used. Models can miss details in the middle of a long input, even when all of it fits. If the file doesn’t fit at all, the tool cuts, summarises or picks parts, and it doesn’t always say which.

Why it matters. This is how you get a fix that uses a variable the AI never saw declared. It’s also how a summary skips the clause on page 40.

Watch out. “Use a tool with a bigger window” is half an answer. A bigger window holds more, and it doesn’t decide what matters. Send the part that matters plus the definitions it depends on.

5. What do temperature and top-p do?

Both control how the model picks the next token. Temperature sets how adventurous the choice is. Top-p keeps only the likeliest options whose chances add up to a set share, such as 90 percent.

Why it matters. For SQL, summaries and data extraction, you want low temperature and repeatable answers. For brainstorming names or test data, a higher setting gives more variety.

What the interviewer wants. That you match the setting to the job, and that you don’t promise identical output. Even at zero temperature, many services don’t guarantee the same answer twice.

Five minutes, any AI assistant

Ask for the same query twice. First with nothing:

Write a SQL query for our top customers.

Then in a fresh chat, with the window filled properly:

Write a SQL query for SQL Server. Table: Sales.Orders (CustomerID, OrderDate, Amount). Return the 10 customers with the highest total Amount in 2024, with their totals.

Count the guesses in the first answer: table names, column names, the date range and what “top” even means. That gap is the whole lesson about context, and it took you a minute.

Why this is worth a book

Chapter 2 has ten questions. It also covers system prompts, what happens when a conversation outgrows the window, zero-shot against few-shot prompting, prompt caching and why longer prompts cost more.

The Senior question in that chapter is the one worth rehearsing: how would you give a model a large schema or codebase without blowing the window?

100 AI Interview Questions and Answers for Data Professionals has 184 key terms alongside the hundred questions. It’s in Kindle (also on Amazon.in), paperback and audiobook.

Part three is about the square on the grid where the money goes: wrong, and sounding sure.

Out of the window is not out of scope, it is out of mind.

Published by Pinal Dave on SQLAuthority. More of my work at pinaldave.com.

First appeared on Tokens and the Context Window: Out of the Window, Out of Mind

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories