Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162511 stories
·
33 followers

Requirements Traceability Matrix for AI Software Delivery—From Issue to Pull Request

1 Share

Your issue tracker and version control platform hold most of the requirements traceability matrix an auditor wants, in a form nobody in your org can query.

On a Tuesday, an auditor asks a simple question. Which requirement produced the code in release 4.2? Who verified that it worked? Instead of answering from Jira and GitHub, someone spends two days reconstructing the story from merged pull requests and old release notes. It’ll still be a best guess, delivered with total confidence.

You need a requirements traceability matrix. It connects each requirement to the work that implements it, then traces every delivered artifact back to the reason it exists. Its purpose is to make the auditor’s question boring.

In most engineering orgs, it does the opposite, because someone has to keep it current by hand. That manual work doesn’t scale, and nobody was ever promoted for updating it. Meanwhile, your development workflow is already creating the links the matrix needs, inside systems you already pay for, where nobody can query them.

What the Matrix Was Supposed to Do


Image generated with AI

The job of the matrix is to connect the reason for a change to the issue and pull request that carried it. It also ties that work to the test run and approval that checked it. It answers two questions. Given a requirement, what proves it shipped? Given a diff, what business reason justifies it?

That first question matters outside an audit. A team that can trace a requirement to the work behind it can understand unfamiliar code faster.

But the matrix usually gets treated as audit paperwork. When that happens, it never earns budget. It runs on whatever time is left over, and that time keeps failing to materialize.

The research says underfunding traceability costs more than it saves. A 2015 controlled experiment gave participants real maintenance tasks on unfamiliar projects. Some tasks included traceability links between requirements and code. Participants who could follow those links finished 24% faster on average and produced 50% more correct solutions.

Faster work frees capacity, and capacity determines what your team can take on. Traceability, it turns out, is a productivity feature wearing a compliance costume.

The Chain Your Workflow Leaves Behind

Set the hand-kept matrix aside and trace a change through your org. Each handoff leaves a record whether anybody wants it to or not:

  • Requirements arrive from customer commitments or regulations.
  • A Jira or Azure DevOps issue names the requirement.
  • The implementation plan lives in the issue description or a linked design doc.
  • Commits on a branch reference the issue by number.
  • CI runs unit tests and acceptance tests against those commits.
  • A reviewer approves or rejects the pull request, on the record, with a timestamp.
  • The pull request merges and, closing keyword satisfied, closes the issue it named.

Requirement to issue to plan to code to tests to review to merge.

LinkUsual sourceProof
RequirementProduct brief or contractWhy the work exists
IssueJira or Azure DevOpsWhat work was approved
PlanIssue description or design docHow the change should be made
Code changeCommit or pull requestWhat changed
Test resultCI log or test reportWhether verification passed
ReviewPull request approvalWho accepted the change

How much of that chain your platforms hold varies by team. One team writes fixes #4127 in every commit. The next puts PROJ-4127 in the pull request title. A third uses only a branch name like adam/auth-timeout-fix, because its tech lead is confident the branch name speaks for itself. It doesn’t.

Those references don’t join unless your teams use the same pattern. GitHub accepts several closing keywords, for example. Picking the one every team uses and enforcing it with a required status check is this quarter’s decision. It’s a boring decision, which is why it keeps not getting made.

Agent Volume Breaks the Hand-Kept Matrix

Now put coding agents into that workflow. The number of branches and pull requests rises, while human memory of why each change happened stays thin. The hand-kept matrix stops getting updated because somebody has to open it and add the link. That somebody has a release to ship, and the matrix loses that argument every single time.

Agent-authored change puts three questions on your desk. Who reviewed the work? Where is the agent’s implementation plan stored? What Git author or pull request author identifies the agent and the human who directed it? A shared bot account answers none of this, though it does look tidy in the commit log.

The honest objection is that this chain proves custody, not correctness. Commit text like fixes #4127 proves only that somebody typed a number. Tests and review still carry the correctness question. GitLab, for instance, removes existing approvals when new commits land on the source branch by default, so an agent push resets them.


Traceability Check: Pick one merged pull request. Working from Jira and GitHub alone, name the requirement it served. Then name the person who approved it. If you have to ask a developer, that record does not exist.


Your CI Retention Expires the Evidence on a Timer

Agent volume is not the only clock running on the evidence chain from issue to merge. If you ship a high-risk AI system into the EU market, Annex IV of the EU AI Act requires technical documentation carrying test logs and reports dated and signed by the responsible persons, and Article 11 requires it stay current.

Staying current under a rule like that assumes the evidence still exists. Test runs and approvals are worth only what your platforms keep. Most hosted CI platforms delete workflow logs and test reports after a default retention period nobody in your org chose.

GitHub is the worked example. It retains workflow artifacts and logs for 90 days by default, adjustable up to 400 days on private repositories. The JUnit report and workflow log for a March release can be gone by June unless somebody changed a setting that has never once come up in a planning meeting.

Months later, your biggest customer’s security team asks for the verification record of one release. The issue and pull request are still there. The test run was deleted 90 days after it passed, and the reviewer who approved the code has left. You’re attesting to work you can no longer show, which is a memorable position to negotiate a renewal from.

Traceability Is a Read Problem

Retention keeps the evidence alive; finding it is a different problem. Most traceability programs fail on the write side, which your workflow already handles. What your org lacks is a read path.

One query should answer: show me requirement REQ-128, the issue that implemented it and the pull request that merged it. Then show the tests that passed and the reviewer who approved it.

Software supply chains already treat machine-emitted records as proof. SLSA provenance makes the build platform record how an artifact was made. Human-written claims about a build stay claims; machine-emitted ones carry their own proof.

Your workflow leaves the same kind of records behind. Building the read path means a scheduled job that calls the Jira and GitHub APIs, then stores issue IDs and approval timestamps.

Raise your CI artifact retention window this week, before you scope that job. On a private repository, that one setting buys 310 more days of evidence, which is the best return on a mouse click you’ll get this quarter.


Explore Progress Forge Orchestration

Turn AI-assisted development into a repeatable engineering process by orchestrating the coding agents you already use through structured, configurable development workflows—from work item to pull request.

Your coding agents execute the work. You own the process. Forge orchestrates it.

Try Forge Now

 

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Getting started with DSC 3.0 – Part 5: Guarding your configuration with Assertion

1 Share

At the end of the previous post I said that dependsOn only decides the order in which resources are processed. It doesn't check that the resource you depend on is actually in the desired state.

Sometimes that's not enough. You want to say "only touch this machine if it's Windows Server 2025", or "only configure the website if IIS is really installed". That's what the Microsoft.DSC/Assertion resource is for.

Note: this post is part of a bigger series on Microsoft Desired State Configuration.

What is an assertion?

An assertion is a group resource. It contains a nested configuration document and runs the test operation on every resource inside it. It never changes anything.

  • If every nested resource is in the desired state, the assertion passes.
  • If one of them isn't, the assertion fails, and DSC doesn't invoke the resources that depend on it.

Only the resources that dependsOn the assertion are affected. The rest of the document runs as usual.

Remark: resources that only implement get, like Microsoft/OSInfo, are a natural fit for assertions. They can read the state but not change it.

A minimal example

Save this as assertion.dsc.yaml:

$schema: https://aka.ms/dsc/schemas/v3/bundled/config/document.json
resources:
- name: Demo key
  type: Microsoft.Windows/Registry
  properties:
    keyPath: HKCU\Software\Demo
    _exist: true
  dependsOn:
  - "[resourceId('Microsoft.DSC/Assertion', 'Windows only')]"
- name: Windows only
  type: Microsoft.DSC/Assertion
  properties:
    $schema: https://aka.ms/dsc/schemas/v3/bundled/config/document.json
    resources:
    - name: Operating system
      type: Microsoft/OSInfo
      properties:
        family: Windows

Two things to notice:

  • The assertion has its own $schema and resources, nested inside properties
  • The registry key depends on the assertion with the same resourceId() syntax we used in Part 4, with Microsoft.DSC/Assertion as the type and the instance name

Apply it:

dsc config set --file assertion.dsc.yaml

The operating system is Windows, so the assertion passes and DSC creates the key. 

Check it:

Test-Path HKCU:\Software\Demo


Let it fail

Remove the key first:

Remove-Item HKCU:\Software\Demo

Now change family: Windows to family: Linux in the document and run the same command again:

dsc config set --file assertion.dsc.yaml

The assertion fails, so DSC doesn't invoke the registry resource. Run Test-Path again and the key is not there.

Remark: the registry resource is listed before the assertion in the document. Like we saw in Part 4, the order in the document doesn't matter. The dependency does.

A more realistic case

A common use is applying settings depending on the Windows Server version. The OS version check is the gate:

$schema: https://aka.ms/dsc/schemas/v3/bundled/config/document.json
resources:
- name: Baseline tag
  type: Microsoft.Windows/Registry
  properties:
    keyPath: HKLM\SOFTWARE\Contoso
    valueName: Baseline
    valueData:
      String: Server2025
  dependsOn:
  - "[resourceId('Microsoft.DSC/Assertion', 'Server 2025 check')]"
- name: Server 2025 check
  type: Microsoft.DSC/Assertion
  properties:
    $schema: https://aka.ms/dsc/schemas/v3/bundled/config/document.json
    resources:
    - name: Server 2025
      type: Microsoft/OSInfo
      properties:
        version: "10.0.26100"

Add one pair like this per OS version and each machine only gets the settings that fit.

Here the version is an exact match. DSC 3.3 added version comparison to Microsoft/OSInfo, so you can assert a constraint instead. 

Remark: this one writes to HKLM, so run it from an elevated terminal.

Back to our previous example

In Part 4, dependsOn made sure IIS was installed before the web resources. But dsc config test on a clean machine could still fail, because the IIS cmdlets weren't there yet.

An assertion that checks the Web Server role fixes that. Add this to the document from Part 4:

- name: IIS is installed
  type: Microsoft.DSC/Assertion
  properties:
    $schema: https://aka.ms/dsc/schemas/v3/bundled/config/document.json
    resources:
    - name: Web server role
      type: PSDesiredStateConfiguration/WindowsFeature
      directives:
        requireAdapter: Microsoft.Adapter/WindowsPowerShell
      properties:
        Name: Web-Server
        Ensure: Present
  dependsOn:
  - "[resourceId('PSDesiredStateConfiguration/WindowsFeature', 'IIS')]"

And let the web resources depend on the assertion. For the application pool:

  dependsOn:
  - "[resourceId('Microsoft.DSC/Assertion', 'IIS is installed')]"

Give the website and the web application the same dependency, next to the ones they already have.

The flow is now: install IIS, check that IIS is installed, and only then configure the pool, the site and the application. If the check fails, DSC doesn't touch them.

Things to keep in mind

  • An assertion only gates the resources that depend on it
  • It never changes anything. If you want DSC to fix something, that's a normal resource
  • The assertion itself must be a top-level instance. A resource inside a group can only depend on its neighbors in the same group, so the IIS install and the assertion stay at the top

We are not there yet. In our next and probably last post about DSC, we'll add AI into the mix. 

Read the whole story
alvinashcraft
16 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

How can undefined opcodes ud0 and ud1 have parameters? How undefined were they?

1 Share

Gunnar Dalsnes wondered how ud0 and ud1 could have parameters if they were undefined? “Does it mean they were not completely undefined, just undocumented and not completely implemented?”

The opcodes ud0 and ud1 have no definition, but that’s not the same as being architecturally an “undefined instruction”. They live in a purgatory where they were not assigned a meaning, but were also not officially declared to be meaningless.

Originally, these byte sequences went through the instruction decoder and happened to slip through a few cracks before somebody finally noticed, “Wait a second, I don’t know how to execute this.”

You can see this when you look at the instructions that are encoded as 00001111 11111xxx, as of the Pentium III.

Bits Bytes Opcode Operand 1 Operand 2 Meaning
00001111 11111000 0F F8 PSUBB mm mm/m64 Subtract packed bytes
00001111 11111001 0F F9 PSUBW mm mm/m64 Subtract packed words
00001111 11111010 0F FA PSUBD mm mm/m64 Subtract packed dwords
00001111 11111011 0F FB No meaning assigned
00001111 11111100 0F FC PADDB mm mm/m64 Add packed bytes
00001111 11111101 0F FD PADDW mm mm/m64 Add packed words
00001111 11111110 0F FE PADDD mm mm/m64 Add packed dwords
00001111 11111111 0F FF No meaning assigned

The byte sequences 0F FB and 0F FF had yet to be assigned a meaning. You can see how the instruction decoder could take a shortcut and say, “Well, all the instructions in this range, or at least all the ones that I care about, take an mm registers and an mm/m64 operand, so I’ll just save myself some transistors and decode all of them with two parameters (mm, mm/m64).” And then after decoding, it would use bit 3 to decide whether to set up the arithmetic unit for an add or subtract, and it would use bits 0 and 1 to decide how to subdivide the bits into saturating units.

And if you gave it a 0F FF, it would be only in that last step that the decoder would realize “Oh dear, I don’t know what to do with a bit combination of 11. I’ll raise an invalid opcode instruction.”

The invalid opcode instruction got raised after the operands were parsed.

You can see the trouble that 0F FF created when those empty slots started to get filled in by the SSE instructions.

Bits Bytes Opcode Operand 1 Operand 2 Meaning
00001111 11111000 0F F8 PSUBB mm mm/m64 Subtract packed bytes
00001111 11111001 0F F9 PSUBW mm mm/m64 Subtract packed words
00001111 11111010 0F FA PSUBD mm mm/m64 Subtract packed dwords
00001111 11111011 0F FB PSUBQ mm mm/m64 Subtract packed qwords
00001111 11111100 0F FC PADDB mm mm/m64 Add packed bytes
00001111 11111101 0F FD PADDW mm mm/m64 Add packed words
00001111 11111110 0F FE PADDD mm mm/m64 Add packed dwords
00001111 11111111 0F FF I want to put PADDQ here but I can’t

the natural place to put the PADDQ instruction is 0F FF, but people had already been using 0F FF with the expectation that it raises an illegal instruction exception. Making it a valid instruction would break those programs, so Intel had to move PADDQ to the rather awkward location 0F D4.

Bonus chatter: Undefined instructions with parameters are actually not uncommon. For example, on AArch64, there is a range of 65,536 instructions set aside as permanently undefined, so the udf instruction takes a 16-bit immediate to specify which invalid opcode you want. The PDP-10 reserved opcode 000 as a permanently illegal instruction, and it carries a register and a memory address as parameters. (Because all PDP-10 instructions carry a register and a memory address as parameters.)

There are also so-called “unofficial opcodes” which are instructions that are not part of the instruction set architecture, but for which people reverse-engineered a consistent behavior and began to rely on it. (The 6502 processor is well-known in nerd circles for having undergone this type of analysis.) The 0F FF is one of these “unofficial opcodes” that was popular enough that Intel felt pressure to maintain backward compatibility with it, even though it was never architecturally documented or supported.

Bonus bonus chatter: It appears that the mnemonic ud1 was introduced by the nasm assembler:

* Added the following new instructions: SYSENTER, SYSEXIT, FXSAVE,
  FXRSTOR, UD1, UD2 (the latter two are two opcodes that Intel
  guarantee will never be used; one of them is documented as UD2 in
  Intel documentation, the other one just as "Undefined Opcode" --
  calling it UD1 seemed to make sense.)

It seems obvious that ud1 is also the name that Intel gave internally to that legacy instruction. Otherwise, there would be no need to call the new one ud2!

The post How can undefined opcodes <CODE>ud0</CODE> and <CODE>ud1</CODE> have parameters? How undefined were they? appeared first on The Old New Thing.

Read the whole story
alvinashcraft
22 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Tokens and the Context Window: Out of the Window, Out of Mind

1 Share

Part two of ten. Most of the surprises people hit with AI at work trace back to this one chapter. A model reads in tokens, holds only so many at once, and charges for each one.

Figure showing that the model sees only what fits in its context window, and everything else is gone.

1. What is a token?

A token is the unit of text a model reads and writes. It can be a whole word, part of a word, a chunk of a number or a punctuation mark. For common English text, a rough rule is about four characters per token.

Why it matters. Tokens are how models are measured and billed. Context limits, prices and speed are all counted in tokens. Code and SQL use more tokens than you’d guess, because a name like SalesOrderDetailID splits into several pieces.

Watch out. The four-characters rule is a rough guide for English. Other languages, numbers and code can run much higher. When the number matters, use the vendor’s own token counter.

2. Why does AI struggle to count letters?

Because a token isn’t a word. A short, common word is usually one token, while a long or rare word splits into several. The model sees a few tokens, not ten letters, which is why it can miscount the letters in strawberry.

Why it matters. The same blind spot shows up in data work. Character counts, string lengths and exact spelling checks are weak spots. Let SQL or a script do the counting, and let the model explain the result.

The follow-up. “Why does it get simple arithmetic wrong sometimes?” Numbers are split into tokens too, and the model predicts digits rather than calculating them.

3. What is a context window?

The most text a model can consider at once, measured in tokens. It holds the instructions, the conversation so far, any files you attached and the answer being written. Anything outside the window doesn’t exist for the model.

What the interviewer wants. That on most models the answer counts against the same window. A long input can leave too little room for the reply.

Watch out. Window sizes change with every model release, so don’t memorise a number. Say how you’d find it: the model’s documentation lists it.

4. Why does AI lose track of a long file?

Fitting in the window doesn’t mean every detail gets used. Models can miss details in the middle of a long input, even when all of it fits. If the file doesn’t fit at all, the tool cuts, summarises or picks parts, and it doesn’t always say which.

Why it matters. This is how you get a fix that uses a variable the AI never saw declared. It’s also how a summary skips the clause on page 40.

Watch out. “Use a tool with a bigger window” is half an answer. A bigger window holds more, and it doesn’t decide what matters. Send the part that matters plus the definitions it depends on.

5. What do temperature and top-p do?

Both control how the model picks the next token. Temperature sets how adventurous the choice is. Top-p keeps only the likeliest options whose chances add up to a set share, such as 90 percent.

Why it matters. For SQL, summaries and data extraction, you want low temperature and repeatable answers. For brainstorming names or test data, a higher setting gives more variety.

What the interviewer wants. That you match the setting to the job, and that you don’t promise identical output. Even at zero temperature, many services don’t guarantee the same answer twice.

Five minutes, any AI assistant

Ask for the same query twice. First with nothing:

Write a SQL query for our top customers.

Then in a fresh chat, with the window filled properly:

Write a SQL query for SQL Server. Table: Sales.Orders (CustomerID, OrderDate, Amount). Return the 10 customers with the highest total Amount in 2024, with their totals.

Count the guesses in the first answer: table names, column names, the date range and what “top” even means. That gap is the whole lesson about context, and it took you a minute.

Why this is worth a book

Chapter 2 has ten questions. It also covers system prompts, what happens when a conversation outgrows the window, zero-shot against few-shot prompting, prompt caching and why longer prompts cost more.

The Senior question in that chapter is the one worth rehearsing: how would you give a model a large schema or codebase without blowing the window?

100 AI Interview Questions and Answers for Data Professionals has 184 key terms alongside the hundred questions. It’s in Kindle (also on Amazon.in), paperback and audiobook.

Part three is about the square on the grid where the money goes: wrong, and sounding sure.

Out of the window is not out of scope, it is out of mind.

Published by Pinal Dave on SQLAuthority. More of my work at pinaldave.com.

First appeared on Tokens and the Context Window: Out of the Window, Out of Mind

Read the whole story
alvinashcraft
35 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Build and Train a 25M Parameter LLM From Scratch on Your CPU

1 Share

You don't need a massive GPU cluster to learn how modern frontier AI works. In our new video on the freeCodeCamp.org YouTube channel, you will build, pre-train, and fine-tune a working 25-million parameter language model directly on your laptop or standard CPU.

Frontier labs train massive systems with billions of parameters, but the core architecture, math, and workflows are fundamentally identical. Scaling the architecture down to 25 million parameters with a compact byte-level vocabulary strips away expensive cloud compute costs and lets training loops run locally on your CPU in seconds. This rapid iteration allows you to directly observe how changes to data curricula, loss functions, and reward designs alter model behavior in real time.

Here are some things covered in the course:

  • Modern LLM Architecture
    Learn how cutting-edge techniques work under the hood, including hybrid linear/sparse attention, Mixture of Experts (MoE), tied embedding weights, and multimodal image patch inputs.

  • Pre-Training from Zero
    Watch the model start from generating random characters and learn Python syntax and structure step-by-step using cross-entropy loss.

  • Post-Training with Reinforcement Learning (RL)
    Build a lightweight RL loop where the model writes code, runs against automated unit tests, and receives rewards—boosting its task success rate from 0% up to passing the majority of evaluations.

  • How AI Researchers Actually Work
    Learn how to think like a modern researcher by forming testable hypotheses, tweaking temperatures and reward structures, and checking for real statistical significance across multiple seeds.

Watch the full tutorial for free on the freeCodeCamp.org YouTube channel (1-hour watch).



Read the whole story
alvinashcraft
45 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

IDE vs CLI: Mastering Efficiency in Modern DevOps Workflows

1 Share

DevOps engineers are increasingly relying on powerful tools to streamline their workflows. Two such critical tools that have become indispensable are the Integrated Development Environment (IDE) and the Command-Line Interface (CLI). While each serves unique purposes, understanding their complementary roles in a DevOps setting is crucial for efficient software development and deployment.

The Dual Roles of IDEs and CLIs

Based on content from IBM Technology

Let’s dive into how IDEs and CLIs serve different needs but often intersect in the modern DevOps workflow. An IDE is primarily your go-to environment for writing and understanding code. It provides features that help you visualize your codebase, refactor it for better readability, and debug with ease. An IDE excels at giving developers a comprehensive view of their projects, enhancing comprehension and speed.

On the flip side, a CLI is your powerhouse for operating systems. It allows for automation of tasks using scripts, management of complex pipelines, and control of cloud deployments. Whether it’s executing test suites or managing configurations with tools like Ansible, the CLI is the choice for operations that require scriptability and remote execution capabilities.

The Core Challenge: Context Switching

Despite their benefits, working with both an IDE and a CLI involves frequent context switching, which can become a productivity bottleneck. Developers often find themselves toggling between coding, running tests, deploying builds, and managing cloud resources. Each transfer requires carrying context manually, such as remembering why a change was made or which environment needs updates next.

As Cedric Clyburn and Legare Kerrison discuss, modern development requires seamless coordination not just in coding but across repositories, tests, build pipelines, deployments, and cloud management. The friction lies not in the tools themselves but in maintaining workflow coherence across diverse environments and tasks.

Achieving Synergy: IDEs and CLIs Working in Harmony

Smoothing out the context transitions between IDEs and CLIs could substantially alleviate workflow friction. An ideal integrated system would allow:

  • Multi-step Execution: Handling complex sequences of tasks routinely performed in software development.
  • Coordinated Execution: Employing AI agents to perform continuous, asynchronous tasks without manual intervention.
  • Automated Validation: Instantly checking and verifying changes to minimize human error and enhance reliability.
  • Shell and Remote Access: Enabling comprehensive management of systems beyond local machines.

As AI development tools continue to evolve, there is a huge opportunity to address context overload by integrating these steps more fluidly. This isn’t merely about generating and writing code efficiently, but about bridging those contextual gaps that currently demand human intervention.

Conclusion: Navigating the Future of DevOps

In conclusion, the journey of a DevOps engineer involves leveraging both IDEs for code comprehension and CLIs for system execution. While current workflows demand a balance between these tools, the real focus should be on enhancing cross-tool compatibility to reduce context overload. Innovation in this area could redefine productivity metrics in DevOps, enabling engineers to focus more on creative problem-solving and less on routine handovers.

In this pursuit, as AI continues to empower these tools, engineers might finally strike that long-sought balance, rendering the complex dance between IDEs and CLIs an orchestrated masterpiece.

Read the whole story
alvinashcraft
57 seconds ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories