Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162298 stories
·
33 followers

How to Build Reliable AI Agent Systems for Production

1 Share

The following article was originally published on the Agentic AI Foundation blog site and is being republished here with the author’s permission.

A customer asks your company’s AI agent to update the shipping address on account 123. The agent relies on a support ticket with a typo and updates account 132 instead. The customer relationship management system reports that the update “succeeded.”

The API successfully completed the authorized request, but because it had no way to know which account the customer approved versus what was intended, it saw no error.

Teams often try to improve reliability by tuning the prompt or choosing a stronger model. But the model did exactly what the surrounding system allowed because the system treated the agent’s decision as the final authority. Prompt tuning and stronger models can improve the agent’s responses, but can’t provide guarantees for actions.

Put each decision in a layer that can enforce it

The model should propose an action, while a policy service decides whether the action is allowed. The execution layer can then perform the action within those limits.

Now consider that same agent, but with a policy service in place. The agent would propose updating an account after reading a support ticket. Before the update runs, a policy service can compare the request with the change the user approved. If the request falls outside those limits, the service can reject it even when the agent sounds confident.

After the policy check, the system can save an audit record that connects the request to the result and identifies the policy version used for the decision.

This separation gives each part of the system a job it can handle. The model interprets the request, while deterministic software enforces the conditions that must always hold.

However, policy enforcement depends on the information used to request an action. A policy check can’t correct a decision that was built on context the system should never have trusted.

Treat agent context as untrusted input

Agents receive instructions from more places than the user’s current prompt. A coding agent may read a SKILL.md file or persisted memory from a previous run. Each source can influence what the agent does next.

Stored instructions are useful, but their source may be stale or malicious. Another agent may have written the memory. A downloaded skill may contain instructions that expose credentials.

Therefore, stored context needs provenance. The system should record who wrote the context and when, then verify that linked files still match the versions the writer used. A risk label can also tell the runtime whether a piece of context may guide an answer or authorize a write.

Warnings help, but they don’t remove the risk. In controlled testing described in “Trustworthy Context Is Untrusted By Default,” Shub Argha reports that a trust preamble reduced one class of context contamination from 88.8 percent to 33.3 percent. The remaining failures show why a warning should sit alongside technical checks.

Once a team can tell where context came from, the next decision is what an agent may do with it.

Give the agent only the authority required for the task

OAuth scopes provide a useful boundary, but a broad write scope still leaves a large decision to the agent. The token may allow the agent to update any record even though the user approved one field on one account, as we saw in the opening example.

A narrower capability can represent the exact action the user approved. For example, the runtime could issue a capability that permits one update to the shipping address on account 123. The capability expires after the update, so the agent cannot reuse it for another customer.

Narrow capabilities also apply to local execution. An agent that needs to format one file doesn’t need unrestricted shell access. The runtime can provide a tool with the necessary input and keep other commands unavailable.

As a result, a prompt injection has less authority to work with. The agent may still request the wrong action, but the runtime can reject anything outside the capability it received.

Narrow permissions also make audit records easier to understand because the record contains the authority granted for that specific action. When authority lives in a broad token or a long prompt, an operator has to reconstruct what the agent was supposed to do after the failure.

Even narrow authority doesn’t require an agent to act. A reliable system also needs clear conditions for when the agent should stand down.

Make stopping part of normal operation

An agent can cause damage even without calling a sensitive tool. It can post a wrong answer to a customer or keep replying after a human has taken over.

The runtime should make restraint part of the workflow. If the agent’s confidence falls below a set threshold, the runtime can route the support ticket to a human representative. A reply from the human can then cancel any pending response from the agent.

Confidence checks and rules that stop the agent when a human takes over belong in the routing and execution layers. The model shouldn’t make either decision on its own. Similarly, a kill switch must stop an agent even when the model is in the middle of a plan.

Stopping safely solves one part of reliability, but long-running agents also need a plan for failures that occur after valid work begins.

Preserve state so the system can recover

Consider a browser agent that has filled out most of a form when the page changes. A retry that starts from the beginning could submit an earlier step twice, while a retry that guesses where to continue may skip a required field.

The execution system should record each confirmed step and attach an idempotency key to any action that must happen once. The key is a unique identifier that tells the server a retry belongs to the same action. Then, when the workflow resumes, the system can continue from the last confirmed state without repeating a completed action.

Agent sandboxes create a related problem. Keeping every sandbox running during long idle periods wastes resources and keeps execution environments available longer than needed. Hibernation can reduce both concerns, but only when the wake-up process restores the required state and handles a failed resume.

Recovery becomes much harder when the system can’t explain what happened before the interruption.

Keep evidence that people and agents can inspect

An execution log should connect the context an agent received to the action it requested. It should also show which policy allowed the action and what result came back.

Teams can use that record during an incident, but the agent can also use it during later work. For example, an agent that can read the reason behind a previous code change doesn’t need to infer intent from the final diff alone.

Context graphs offer one way to preserve those relationships across time. A graph can connect a tool call to the policy that governed it. The same record can link the source context to the outcome. Later, a person or agent can query the relationships to understand why the system made a decision.

The record still needs limits. Secrets shouldn’t be copied into an audit trail, for example, and retention rules should match the data involved.

A deliberate record is safer than relying on a conversation transcript and hoping it contains everything an operator will need.

A reliable agent system can explain why an action was allowed and resume safely when that action fails. Enforceable policies and recorded state provide those guarantees around the model.

Keep the controls portable

Many organizations will use more than one model or agent runtime. When each agent carries its permissions inside a prompt, the rules can drift as teams add models and tools.

MCP gives clients and servers a shared way to describe and call tools. Teams can use that common tool surface to enforce authorization at the server or gateway, regardless of which model requested the action.

A shared context format can also preserve provenance when a workflow moves between agents. Open specifications provide consistent interfaces, while each deployment remains responsible for its policies and enforcement.



Read the whole story
alvinashcraft
43 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

When Agile Team Trust Breaks Before the Scrum Master Acts | Oluyomi Emmanuel

1 Share

Oluyomi Emmanuel: When Agile Team Trust Breaks Before the Scrum Master Acts

Read the full Show Notes and search through the world's largest audio library on Agile and Scrum directly on the Scrum Master Toolbox Podcast website: http://bit.ly/SMTP_ShowNotes.

 

"Once that trust is broken, everything else, nothing else matters." - Oluyomi Emmanuel

 

Oluyomi Emmanuel started his Scrum Master journey after a strategic move toward Agile ways of working changed his project coordinator role into something deeper: helping people communicate when the answers were unclear. In this episode, he shares a hard lesson from a newly formed software team of eight people, where tension between a tech lead and a new developer first appeared in GitHub pull request comments. There were sarcastic replies, cameras switched off during the daily standup, and a growing distance when the team worked in the same office. Oluyomi initially treated it as a normal forming-stage conflict and waited for the team environment to absorb the tension. By the time the tech lead escalated concerns to the developer's line manager, trust had already deteriorated. A facilitated feedback session produced an improvement plan, but the developer experienced the feedback as judgment rather than support and left the team less than a month later. Oluyomi's lesson was clear: a Scrum Master needs to name trust concerns early, first in one-on-ones, then together, before the relationship becomes too damaged to repair.

 

In this episode, we refer to psychological safety, conflict resolution, and team trust.

 

Self-reflection Question: What early sign of broken trust have you noticed but not yet named with your team?

 

[The Scrum Master Toolbox Podcast Recommends]

🔥In the ruthless world of fintech, success isn't just about innovation—it's about coaching!🔥

Angela thought she was just there to coach a team. But now, she's caught in the middle of a corporate espionage drama that could make or break the future of digital banking. Can she help the team regain their mojo and outwit their rivals, or will the competition crush their ambitions? As alliances shift and the pressure builds, one thing becomes clear: this isn't just about the product—it's about the people.

 

🚨 Will Angela's coaching be enough? Find out in Shift: From Product to People—the gripping story of high-stakes innovation and corporate intrigue.

 

Buy Now on Amazon

 

[The Scrum Master Toolbox Podcast Recommends]

 

About Oluyomi Emmanuel

 

Oluyomi is a Scrum Master, Agile Coach, and Product Lead with over seven years of experience helping teams work better and deliver meaningful digital products. Having worked across energy, consumer platforms, and fintech, he is passionate about building high-performing teams, creating clarity, and turning Agile principles into real business and customer value.

 

You can link with Oluyomi Emmanuel on LinkedIn.

 





Download audio: https://traffic.libsyn.com/secure/scrummastertoolbox/20261005_Oluyomi_Emmanuel_M.mp3?dest-id=246429
Read the whole story
alvinashcraft
44 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

137. New Betatalks the Podcast Episodes Coming - with Rick & Oscar

1 Share

New episodes coming, stay tuned!

About Betatalks: watch our podcast videos and follow us on Instagram and LinkedIn. 





Download audio: https://www.buzzsprout.com/1622272/episodes/19898817-137-new-betatalks-the-podcast-episodes-coming-with-rick-oscar.mp3
Read the whole story
alvinashcraft
44 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Sam Nasr: AI Transformation - Episode 422

1 Share

https://clearmeasure.com/developers/forums/

Sam Nasr is a Senior Software Engineer and Trainer at NIS Technologies in Cleveland, Ohio, specializing in the Microsoft AI stack, Azure, and .NET development. A software developer since 1995, Sam is an 8-time Microsoft MVP and Microsoft Certified Trainer who leads the Cleveland C# User Group, the .NET Study Group, and the Azure Cleveland User Group. He shares his expertise as a contributing author for Visual Studio Magazine and through his LinkedIn Learning courses, and he blogs regularly at samnasr.blogspot.com, This will be his second appearance on the podcast, having previously joined me to discuss SQL Server for developers.

Sam's Blog - https://samnasr.blogspot.com/
LinkedIn - https://www.linkedin.com/in/samsnasr/
Github - https://github.com/samnasr
X Account - https://x.com/samnasr
Upcoming Event - https://www.meetup.com/caparea-net/events/316713608
Linktree - linktree.com/samnasr

Previous Appearances on the Azure & DevOps Podcast:
Episode 122 - https://azuredevopspodcast.clear-measure.com/sam-nasr-on-sql-server-for-developers-episode-122

Want to Learn More?
Visit AzureDevOps.Show for show notes and additional episodes.





Download audio: https://traffic.libsyn.com/clean/secure/azuredevops/Episode_422.mp3?dest-id=768873
Read the whole story
alvinashcraft
44 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Seeing, understanding, and responding: low-power CNN, VLM and SLM workloads on Raspberry Pi 5

1 Share

Most edge AI/ML workloads belong on the CPU, but accelerators are important for the small minority of workloads that the CPU can’t accommodate. Raspberry Pi users have a variety of choices for acceleration, with our own AI HAT+ and HAT+ 2 as well as offerings from the wider Raspberry Pi ecosystem. Here, our friends at Sixfab and DEEPX discuss running vision and compact language AI on Raspberry Pi 5 with the Sixfab AI HAT+, and what a three-watt NPU makes possible at the edge.

Raspberry Pi has made AI development accessible to millions of engineers, students and makers. The next step is moving from a model that runs once, in a demo, to an intelligent system that keeps watching, understanding and responding — without depending on a constant cloud connection.

The Sixfab AI HAT+ for Raspberry Pi 5 is a third-party AI accelerator board built around the DEEPX DX-M1M NPU. It adds 25 TOPS of dedicated AI acceleration at approximately 3 W of typical sustained NPU power, while leaving the Raspberry Pi’s CPU free for camera handling, application logic, connectivity and device control. In this article we’d like to show what that combination is genuinely good at, where it isn’t the right tool, and how to get from first boot to first inference in a few minutes.

Why three watts, not the TOPS number, is the headline

Peak TOPS figures make good marketing, but most edge products live and die by their power and thermal budget. The real system still needs headroom for the Raspberry Pi itself, a camera, storage, networking and whatever the device actually controls. An accelerator that stays around 3 W under sustained load means smaller enclosures, simpler (often passive) cooling, and AI that can run continuously rather than in short bursts.

That efficiency is what makes local AI practical in the places where it matters most: smart cameras, robots, industrial equipment and distributed sensors, where cloud latency, connectivity or data privacy would otherwise become the limiting constraint.

ENGINEERING NOTE:  The ~3 W figure refers to typical sustained NPU power for the 25 TOPS DX-M1M option, not complete system wall power. Always measure your final Raspberry Pi configuration with its actual camera, storage, workload and cooling.

One NPU for three classes of intelligence

Most edge products begin with perception: a CNN detects an object, segments a work area, estimates a pose or flags an abnormal event. The next generation of edge systems also needs to understand context and communicate what it sees. The DX-M1M is designed to support that whole journey on a single piece of silicon and one toolchain.

CNN: efficient visual perception

Object detection, classification, segmentation, pose estimation, depth and image enhancement turn camera pixels into structured events. The DEEPX ModelZoo ships pre-optimised examples — including YOLO-family detectors, segmentation and pose models — and DX-Stream helps you assemble them into GStreamer-based camera pipelines.

Compact VLM: connecting vision with meaning

A compact vision-language model adds contextual understanding on top of visual results: describing a scene, answering a focused question about an image, or interpreting an event in natural language — locally, with nothing leaving the device.

SLM: talking to the edge system

A small language model can interpret a user command, summarise local events or generate a concise explanation, creating a natural interface for cameras, robots and machines without routing every interaction through the cloud.

What about memory?

The first question any developer asks about language models at the edge is how much memory is available and how large a model will fit. The DX-M1M module carries 2 GB of dedicated on-module LPDDR4x, which comfortably supports VLM and SLM workloads alongside vision models.

Accuracy matters more than a peak benchmark

Edge AI performance is only useful if the model stays accurate on the images and conditions your application actually faces. DEEPX treats INT8 optimisation as an accuracy-aware engineering process, not a simple conversion step:

  1. Start with a known FP32 accuracy baseline and a clearly defined application metric.
  2. Calibrate with representative data from the real deployment environment — not only clean sample images.
  3. Compile and optimise the model with DX-COM into a portable .dxnn deployment artifact.
  4. Compare accuracy, latency, power and thermal behaviour on the final Raspberry Pi system.

This matters most in the difficult conditions edge cameras actually meet: low light, glare, motion blur, occlusion and unusual angles. A credible result shows both the efficiency gain and the retained application accuracy.

How fast is it, really?

Numbers beat adjectives. The figures below were measured on a Raspberry Pi 5 (8 GB) running the shipping Sixfab software release, comparing the DX-M1M and DX-M1 NPUs:

WorkloadDX-M1M (NPU)DX-M1 (NPU)
mobilenet_v2, 240×2402361 FPS3223 FPS
deeplabv3plus, 512×512155 FPS231 FPS
Qwen3-1.7B, 96 prefill tokensTTFT: 599.04 ms TPS: 4.64 tok/sTTFT: 544.58 ms TPS: 11.60 tok/s

* The DX-M1 uses LPDDR5, while the DX-M1M uses LPDDR4X, resulting in a performance difference due to their memory bandwidths.

From first boot to first inference

The fastest way to understand the platform is to verify the hardware, run a packaged model and watch the NPU work. On Raspberry Pi OS:

sudo apt update && sudo apt install apt-repo-sixfab
sudo apt update && sudo apt install sixfab-dx
dxrt-cli -s        # confirm device and software status
run_hello_world    # run a packaged example
dxtop              # observe the NPU in real time

From there, pick an optimised ModelZoo workload or bring your own supported ONNX model through DX-COM. The resulting .dxnn file runs through DX-RT’s C++ or Python APIs, while your Raspberry Pi application keeps handling cameras, business logic, user interfaces, networking and control.

The software stack in one view

ComponentWhat it does for you
DX-COMOptimises and compiles supported ONNX models into DEEPX .dxnn artifacts.
DX-RTRuns .dxnn models through C++ or Python APIs on the target device.
DX-StreamBuilds GStreamer-based capture, inference and post-processing pipelines.
ModelZooPre-optimised examples across detection, classification, segmentation, pose and more.
dxrt-cli / dxtopDevice status, utilisation, temperature, clocks and memory at a glance.

What can you build with it?

Start at home. A camera on your workbench that recognises when your 3D print has failed and messages you a snapshot with a one-line description. A front-door camera that tells the difference between a courier, the neighbour’s cat and someone loitering — and summarises the day’s events each evening, entirely locally. A garden or pet monitor that only alerts you about things worth seeing. All of these follow the same loop: a CNN sees the event, a compact VLM adds context, an SLM explains it or interprets your next command, and your Raspberry Pi application decides and acts.

The same loop scales to serious deployments: continuous local detection in smart cameras with reduced cloud dependence; low-power visual quality and safety monitoring beside a production line; perception and compact multimodal intelligence on robots while the Raspberry Pi runs ROS 2 and control logic; and privacy-aware occupancy and behaviour analytics in retail and smart spaces.

What it is not for

Like any edge accelerator, the AI HAT+ is not trying to compete with cloud inference. Applications that need broad world knowledge, very long conversational context or continuous learning will always run better where compute and memory are effectively unconstrained. The DX-M1M is at its best running tightly scoped, always-on intelligence next to a camera or sensor — where privacy, latency, offline operation and a single-digit-watt power budget are the requirements that actually decide the design. Being honest about that boundary is exactly what makes the platform trustworthy inside it.

Community and what’s next

Everything you need to go deeper is public: DX-AllSuite and the ModelZoo are on GitHub, and the DEEPX Developer Portal hosts documentation, guides and support channels where you can ask questions and file issues.

Where to get one

The 25 TOPS Sixfab AI HAT+ for Raspberry Pi 5 is available now from Sixfab, with the 13 TOPS version expected to launch toward the end of 2026. Documentation and the quick-start guide are linked below.

Links

The post Seeing, understanding, and responding: low-power CNN, VLM and SLM workloads on Raspberry Pi 5 appeared first on Raspberry Pi.

Read the whole story
alvinashcraft
44 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Discontinuing Swift Language IDE Support in the Kotlin Multiplatform Plugin

1 Share

Starting with IntelliJ IDEA 2026.3 and Android Studio Rabbit 2 | 2026.2.2, we’re discontinuing Swift language IDE features in the Kotlin Multiplatform (KMP) plugin. This includes editing, syntax highlighting, and navigation within Swift files and between Swift and Kotlin.

Running iOS applications and debugging Kotlin code on iOS will remain fully supported. Kotlin/Native functionality, such as Swift Export and importing Swift packages into Kotlin projects, is not affected. 

The existing Swift functionality will remain available in earlier versions of IntelliJ IDEA, Android Studio, and the Kotlin Multiplatform plugin, though these older versions will not receive further updates.

Background

We originally developed Swift support for AppCode (our since-discontinued iOS IDE) and for working directly with Xcode projects. This technology was later adapted for Fleet’s polyglot capabilities and, most recently, for working with mixed Kotlin–Swift codebases using the Kotlin Multiplatform plugin.

While analyzing feature adoption, we found that active usage was limited and steadily declining. Our telemetry and user research showed that most iOS developers working on KMP projects prefer using Xcode for Swift development, while teams that primarily work in the Kotlin Multiplatform plugin increasingly use Kotlin and Compose Multiplatform.

Maintaining this Swift tooling, which inherited its core architecture from AppCode, requires substantial ongoing effort to stay compatible with continuous changes in Swift, Xcode, and the IntelliJ Platform. While we previously considered expanding Kotlin–Swift IDE support in the plugin, after weighing the necessary engineering investment against actual usage numbers we decided not to continue.

Looking ahead

We do not currently plan to replace these features. Moving forward, we will continue evaluating workflows and focus on providing excellent Kotlin tooling in IntelliJ IDEA and Android Studio and improving the KMP experience for iOS developers.

We would like to thank everyone who used Swift IDE features, especially those who have been with us since the AppCode days, and to all the contributors who put so much dedication into building and refining Swift support over the years.

Yours, 

The Kotlin team

Read the whole story
alvinashcraft
44 minutes ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories