Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
161210 stories
·
33 followers

An updated look for the Raspberry Pi Desktop

1 Share

Today brings you some optional updates to the Raspberry Pi Desktop, which gains some modern features while keeping its simplicity.

What is currently known as the Raspberry Pi Desktop has been around for over a decade now – it started out as a mildly customised version of the stock LXDE desktop, and over the years has slowly evolved into something with almost entirely different underpinnings, based on Wayland and labwc. The look has been refreshed on several occasions, with new icons, themes and desktop backgrounds, to the point where there is very little visually in common between today’s Raspberry Pi Desktop and the original LXDE.

But one thing which has never changed is the overall feel of the desktop, which, with its taskbar, status icons and main menu launcher, is very reminiscent of the state of the art in desktop design from 30 years or so ago, with Windows 95 and Apple System 7. Not without good cause – both were good examples of an intuitive desktop interface, and they have the benefit of being very familiar to most users. The intention has always been to offer a desktop experience which most people can simply use without a learning curve, and I believe our Desktop has offered that.

Back in the day!

Desktop designs have continued to evolve in that time, and last year, I started to think that some of the features which are now commonplace in other desktops might be beneficial in ours. Since then, I have been working on a few ideas for how we might add these features to the desktop – not losing the fundamental simplicity of it, and not forcing change for the sake of change on users who are happy with the way it is, but just providing some more options.

I do have to stress here, before the anticipated deluge of complaints, that these new features are entirely optional; unless you choose to enable them, your desktop will continue to look and feel exactly the way it always has. These are purely intended to give users who want a more up-to-date desktop experience the option to have one if they want it. (So please, don’t immediately leap to the comments and complain that you don’t like the new look and we should never have changed it – if you do, I’ll know you didn’t actually bother reading what I wrote…)

So, assuming you do fancy a change, what do you get?

The icon dock and launcher

The biggest change is the introduction of an icon dock, which can either replace the existing taskbar, or be used in addition to it. Two new widgets have been added which are intended to be used in the dock – one is a graphical application launcher, and the other an icon-based combined quick launcher and task list.

The raspberry icon for the graphical application launcher opens the launcher screen. (This is intended as a replacement for the main menu widget which has been the standard application launcher up until now.)

The new launcher screen

This shows all the application icons which would have been in the main menu. By default, icons are sorted in alphabetical order, but they can be dragged and dropped to rearrange them. Single-clicking any icon launches the application; right-clicking (or a long press on a touchscreen) brings up a context-sensitive menu similar to that shown in the traditional main menu.

The search box can be used to filter the displayed applications, and the cursor keys and ‘enter’ can be used to navigate and select applications from the keyboard.

Right-clicking the raspberry icon and selecting the ‘Configure Icon Menu Widget’ option gives some configuration options for the launcher.

Enabling alphabetical sorting undoes any drag-and-drop rearrangement and restores the default order of icons.

Enabling menu categories makes the launcher hierarchical, with a top-level screen for application category icons which can then be clicked to enter individual launcher screens for each category – this is useful if you have a large number of applications installed. To help with remembering which applications are in which category, turning on composite category icons shows miniature versions of the first four application icons in each category in the top-level screen instead of the category icons themselves. (If menu categories is enabled, then drag-and-drop can no longer be used to rearrange the launchers; instead, the hierarchy and arrangement used for the main menu is used, which can be edited using the Main Menu panel in Control Centre as before.)

The new launcher screen with menu categories enabled, showing application categories

By default, the icons are displayed overlaid on the current desktop background, but in the event that this looks too cluttered, a transparent or opaque overlay colour can be applied.

The new task list

The combined quick launcher and task list can display a set of the most commonly-used application icons, and will by default show the same set of icons as the original quick launcher from the taskbar. New applications can be added by right-clicking an entry in the main menu or the graphical application launcher and choosing ‘Add to Launcher’ from the context-sensitive menu. As for the application launcher, icons in the quick launcher can be reorganised by dragging and dropping.

Clicking an icon launches the application. Running applications – whether launched from the quick launcher or elsewhere – are shown with a superimposed count of the application’s open windows. If an application was already in the quick launcher, this count is superimposed on the existing icon; if the application was not in the launcher, its icon is added to the end and the count is superimposed on the new icon.

Clicking on an icon for a launched application brings all that application’s windows to the foreground. If all the application’s windows are already in the foreground, a new window is opened (if the application supports multiple windows).

Right-clicking an icon brings up a menu which shows all the application’s open windows, allowing them to be individually brought to the foreground, and offers options for maximising or hiding the application’s windows.

Applications can be removed from the quick launcher by right-clicking the icon when the application is not running, and choosing ‘Remove from Launcher’ from the context-sensitive menu.

These two new widgets are intended to offer alternatives to the existing main menu and window list, and while they are intended for use in the new dock, they can also be used in the original taskbar if desired; likewise, the original main menu and window list widgets can be used in the dock.

Customising the dock

The dock itself can be customised in the same way as the taskbar. A Dock page has been added to the Control Centre and offers options for colour, position and icon size.

One often-requested feature which has been added is the ability for the taskbar and the dock to automatically hide when not in use – if the ‘Autohide’ option is enabled for either, it will slide away when the mouse is not over it, and reappear when the mouse is moved to the edge of the screen where it has hidden. The ‘Exclusive’ option controls whether a maximised window covers the taskbar or dock; if exclusive is set for either, maximised windows will not overlap it. This can be useful if you want, for example, the clock to always be visible.

A new Widgets page has been added to the Control Centre which allows plugins to be placed in either taskbar or dock, and arranged as desired.

The dock offers two locations for widgets. The left-hand side of the dock is the dock proper, while the right-hand side is given over to the ‘tray’. This is intended to hold status icons, which are shown as two rows with the icons half the size of those in the main dock.

When icons are added to the tray, they are automatically split across the two rows such that the length of top and bottom is kept as consistent as possible. If more control over the split is required, the ‘tray split’ widget can be added to the tray widgets – this imposes a hard line break between the widgets at this point.

In order for either the taskbar or the dock to be shown, they must contain at least one widget – if either contains no widgets, it is hidden.

Note that the taskbar and the dock can be used simultaneously – for example, if desired, the status icons can be displayed on the taskbar as they are now, with just the new launcher and task list in the dock. (This is how my own desktop is configured.)

I’d encourage you to play about with the Widgets control panel to come up with whatever arrangement of taskbar, dock and widgets works best for you, but in order to make it easy to try the new features, an option to switch between a few preset desktop styles has been added to the Defaults page in the Control Centre.

Setting the style to ‘Dock’ will remove all widgets from the taskbar, add the new launcher and task list to the dock, and add the status icons in the tray. Setting the style to ‘Taskbar / Dock’ will add just the launcher and task list to the dock, while leaving the status icons on the right of the taskbar. Or if you just want the desktop as it has always been, set the style to ‘Taskbar’, which will leave the existing main menu and task list on the left of the taskbar, with the status icons on the right.

Other changes

One other small change which you can see in some of the pictures above is the addition of an analogue clock mode to the clock widget – simply right-click the digital clock, choose ‘Configure Clock Widget’, and turn on the analogue clock mode. The colour of the face and the hands can also be customised.

In addition to the dock changes, there is now the ability to use a more efficient mechanism for drawing the desktop background picture. This has always been drawn by pcmanfm, the file manager application – as this is the file manager, it allows icons to be placed on the desktop. But if desktop icons are not required, significant memory savings can be made by using the very lightweight swaybg program to draw the desktop.

This can be enabled in the Desktop tab of the Control Centre – switching off ‘Active Desktop’ disables the drawing of the desktop by the file manager and instead uses swaybg to display the same picture. (The Wastebasket which is normally displayed on the desktop can then be found in the Places pane in any file manager window.)

We now recommend disabling the active desktop for any platform with less than 2GB of memory – this is checked at boot, and a notification is displayed suggesting that active desktop be disabled. Also, if you choose either of the new dock-based styles for the desktop in the Defaults page, active desktop is disabled by default, but it can be re-enabled from the Control Centre if desired.

Note that the file manager was previously responsible for automounting removable drives; as this would no longer happen if the file manager was not running all the time, this functionality has now been moved into the ejecter plugin, which will now automount and display a notification when a drive is inserted.

A new Shortcuts page has also been added to the Control Centre. This allows all system keyboard shortcuts to be viewed and edited, and new ones to be created if desired. (Note that this only covers general system shortcuts, not those which are assigned by individual applications.)

One final change which has been made is an enhancement to the creation of screenshots. Previously, hitting the PrtScrn key took a shot of the entire screen and saved it to the Pictures folder. Now, hitting the key takes the screenshot and also opens a dialog offering the option to open the captured file in an image editor, or to copy it to the clipboard. (If you always want to do the same thing, just tick the ‘remember my choice’ checkbox before pressing one of the buttons, and the prompt will not be shown in future.)

Holding down Alt when pressing PrtScrn does the same thing, but first displays a selection rectangle allowing you to select an area of the screen to be captured rather than the whole screen.

How to get it

We’ve released a new image today with all these new features added. As I said at the start, these features are all optional – if you install the new image, the default appearance of the desktop will be the same as it has always been, with a taskbar and no dock; to try the dock, just launch the Control Centre, go to the Defaults page, and select either the ‘Taskbar / Dock’ or ‘Dock’ styles.

Similarly, for existing images, simply update using either the updater widget or the usual

sudo apt update
sudo apt full-upgrade

from the terminal. Again, installing the updates will not enable the dock; use the Control Centre as above to try it.

We do hope you enjoy using the new features – as always, do let us know in the comments or the forums how you get on, and what you think!

The post An updated look for the Raspberry Pi Desktop appeared first on Raspberry Pi.

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Multi-Agent AI Systems: A Smarter Way to Detect and Manage Technical Debt

1 Share

TLDR: Multi-agent AI systems use specialized AI agents to analyze repositories across architecture, security, performance, testing, and code quality. By combining multiple perspectives, they help teams identify risks earlier, prioritize improvements, and address technical debt more systematically.

Modern development teams ship features at an incredible pace. AI coding assistants have accelerated development even further, making it possible to generate large amounts of code in minutes.

But as repositories grow, writing code is no longer the hardest part. Maintaining quality across thousands of files, multiple services, and dozens of contributors becomes increasingly difficult.

Most teams still rely on:

  • Manual code reviews,
  • Static analysis tools,
  • Security scans,
  • Periodic architecture reviews, and
  • Technical debt audits.

These practices are valuable, but they often operate as separate checks rather than providing a unified view of repository health.

A performance bottleneck might remain unnoticed until users complain.

  • Architectural issues may only surface when new features become difficult to implement.
  • Security vulnerabilities can go undetected until much later in the development cycle.

What teams need isn’t just faster code generation. They need a better way to understand repository health and prioritize improvements.

This is where multi-agent AI systems can make a meaningful difference.

What is a multi-agent AI system?

Multi-agent AI system working architecture

A multi-agent AI system uses multiple specialized AI agents that collaborate to analyze and improve a repository. Instead of asking one AI assistant to inspect everything, responsibilities are divided among specialized reviewers.

Different agents focus on areas such as:

  • Architecture,
  • Security,
  • Performance,
  • Testing,
  • Code quality, and maintainability.

Together, they can provide broader repository coverage by analyzing different aspects of software quality simultaneously. The actual value depends on how well the agents are scoped, coordinated, and supported with relevant context.

A typical multi-agent system may:

  • Analyze repository structure,
  • Identify risks and bottlenecks,
  • Detect quality issues,
  • Recommend improvements,
  • Prioritize findings, and
  • Track progress over time.

This transforms governance from an occasional review activity into a more systematic engineering process.

When does a multi-agent approach make sense?

Not every repository needs a multi-agent architecture.

A multi-agent approach often makes sense when:

  • The repository is large or rapidly growing.
  • Multiple quality dimensions need evaluation.
  • Different teams own different parts of the system.
  • Audits are performed repeatedly.
  • Findings must be consolidated and prioritized.

A single-agent approach may be better when:

  • The repository is relatively small.
  • Tasks are narrowly scoped.
  • Execution cost is more important than broad coverage.
  • There is limited cross-domain complexity.

The goal is not to replace simpler processes, but to introduce multiple specialized perspectives when the repository’s complexity justifies it.

Why single-agent approaches can struggle with complex repositories

A single AI assistant behaves much like a capable generalist developer. That approach works well when reviewing a component, investigating a bug, or solving a focused problem.

Modern applications, however, rarely contain isolated concerns.

For example:

  • A performance improvement may introduce security considerations.
  • An architectural change can affect testing coverage.
  • A deployment modification may impact reliability.

As complexity grows, teams often benefit from multiple specialized perspectives.

When agents are well-scoped and properly coordinated, a multi-agent approach can provide broader coverage, reduce blind spots, and make findings easier to organize.

The goal isn’t simply finding more issues. It’s helping teams identify the issues that matter most before they become expensive problems.

Single-agent vs. Multi-agent approaches

Aspect Single-agent approach Multi-agent approach
Responsibilities One agent handles multiple concerns Tasks are divided among specialized agents
Coordination Simpler Requires orchestration
Domain coverage Broad but potentially shallow Specialized across domains
Complexity Lower Higher
Resource usage Usually lower Can be higher
Failure handling Fewer coordination points More coordination and failure points
Best suited for Focused tasks Complex problems requiring multiple perspectives

Multi-agent systems are not automatically better. They introduce additional coordination, execution, and management complexity. Their value comes when specialization provides benefits that justify that added complexity.

Multi-agent architectures in practice

Multi-agent systems are commonly implemented using either centralized or decentralized coordination.

Architecture Best for Trade-off
Centralized Consistent coordination and shared decision-making Dependency on central controller
Decentralized Flexible agent-to-agent collaboration More complex coordination

For repository governance scenarios, the choice depends on whether consistency or flexibility is the higher priority.

Specialized AI auditors for code governance

One of the biggest advantages of multi-agent AI systems is specialization. Rather than asking one agent to evaluate everything, different agents focus on different aspects of software quality.

1. Architecture auditor

Evaluates structural integrity, coupling, scalability concerns, and service boundaries.

Example: Detects excessive dependencies between services that make future changes more difficult.

2. Security auditor

Identifies hardcoded secrets, authentication risks, vulnerable dependencies, and common security weaknesses.

Example: Flags API keys committed directly into source control.

3. Performance auditor

Reviews query efficiency, resource usage patterns, response time risks, and optimization opportunities.

Example: Identifies N+1 database query patterns that may increase latency.

4. Testing auditor

Evaluates test coverage, edge-case handling, and testing gaps.

Example: Detects payment processes with insufficient automated testing.

5. Code quality and maintainability auditor

Identifies duplication, readability concerns, documentation gaps, and maintainability issues.

Example: Highlights duplicate business logic appearing across multiple services.

Together, these auditors help teams develop a more holistic understanding of repository health.

Real-world example: Auditing a large SaaS repository

Consider a SaaS platform with:

  • Multiple microservices,
  • Hundreds of monthly pull requests,
  • Several engineering teams, and
  • Growing technical debt.

The team recently adopted AI-assisted development and significantly increased delivery speed. However, repository health became harder to understand.

Sample prioritized findings

Finding Auditor Priority
Exposed credentials in configuration Security Critical
Missing test coverage in payment process Testing High
Expensive database query affecting API response time Performance High
Increasing service coupling Architecture Medium

Individually, these findings are useful. Together, they reveal broader risks that might otherwise be missed.

Instead of receiving disconnected warnings, the engineering team receives a consolidated, prioritized view of repository health.

This is where multi-agent systems provide meaningful value: helping teams move from isolated observations to informed engineering decisions.

From findings to action

Analysis alone does not improve software quality. Teams must still decide what to fix first and how to implement changes.

Refactoring and planning

A planning component can:

  • Consolidate findings,
  • Remove duplicate recommendations,
  • Prioritize issues,
  • Evaluate impact versus effort, and
  • Provide implementation guidance.

Implementation support

Implementation-focused agents may:

  • Suggest refactorings,
  • Generate code recommendations,
  • Track remediation progress, and
  • Assist with follow-up work.

However, implementation should remain subject to developer review and approval, particularly for security-sensitive, architectural, or production-impacting changes.

The role of these systems is to help teams discover and organize improvements, not replace engineering judgment.

Running your first multi-agent audit in Code Studio

Getting started with multi-agent governance in Syncfusion Code Studio is straightforward. The idea is simple: create a set of specialized AI auditors, point them at your repository, and let them analyze different aspects of code quality in parallel.

Step 1: Create your AI agents

Start by creating the agents you want to use in the .codestudio/agents/ directory.

Select the Configure Custom Agents option in the chat interface. Create as many agents as required.

Creating AI agents in Code Studio
Creating AI agents in Code Studio

Common examples include:

  • Security auditor,
  • Architecture auditor,
  • Performance auditor,
  • Testing auditor, and
  • Code quality auditor.

Each agent is defined using a simple Markdown configuration that describes its purpose and responsibilities.

Multi-agent AI systems allow each auditor to focus on a specific area, resulting in more comprehensive and actionable insights than a single general-purpose reviewer.

Step 2: Open your repository in Code Studio

Once your AI auditors are configured, open the repository you want to analyze in Code Studio and verify that your .codestudio/agents/ directory contains the agents you’ve created.

A typical structure might look like:

.codestudio/
└─ agents/
├─ security-auditor.md
├─ architecture-auditor.md

├─ performance-auditor.md

└─ testing-auditor.md

This directory acts as the discovery point for your audit configuration.

When an audit begins, Code Studio automatically identifies the available agents and invokes them based on their defined responsibilities. By organizing agents this way, you can easily customize the analysis for different repositories.

For example, a security-focused project might use additional security auditors, while a large enterprise application may include agents for architecture, maintainability, and DevOps reviews.

Step 3: Run a repository-wide audit

With your agents configured and repository ready, it’s time to let them get to work.

Run the following prompt in the Code Studio chat window:

Run full audit
Run a repository-wide audit in Code Studio
Run a repository-wide audit in Code Studio

For larger repositories, you can speed up the analysis by running agents in parallel:

Run full audit --parallel

This launches all configured auditors and begins analyzing your repository across multiple areas, including architecture, security, performance, testing, maintainability, and code quality.

Rather than reviewing the codebase from a single perspective, each specialized agent focuses on its own domain and reports findings independently. This allows teams to uncover a broader range of issues in a single audit run, from performance bottlenecks and security risks to architectural concerns and testing gaps.

Step 4: Review the generated reports

After the audit is completed, Code Studio organizes findings into structured reports that help you move from issue discovery to action.

You’ll typically find reports such as:

  • summary/handoff.md → Overall repository health scorecard and key findings.

    summary/handoff.md file

  • summary/<domain>/audit.md → Detailed analysis and identified issues.

    summary/<domain>/audit.md file

  • summary/<domain>/plan.md → Recommended fixes and improvement priorities.

    summary/<domain>/plan.md file

  • summary/<domain>/progress.md → Implementation status and tracking.

    summary/<domain>/progress.md file

Rather than presenting a long list of isolated warnings, these reports provide context around each finding, including where it exists, why it matters, and what steps should be taken next.

Step 5: Prioritize and implement improvements

Finding issues is valuable, but the real impact comes from addressing the right ones first.

Use the generated plans to prioritize improvements based on risk, business impact, and implementation effort. Common starting points include:

  • Fixing critical security vulnerabilities,
  • Addressing architectural bottlenecks,
  • Improving test coverage in high-risk areas, and
  • Resolving performance and scalability concerns.

A key advantage of multi-agent audits is that they help teams focus on what matters most instead of working through a long, unstructured backlog. By highlighting the highest-impact issues first, engineering teams can make measurable improvements to repository health while reducing long-term technical debt.

With a prioritized plan in place, developers can move confidently from analysis to implementation, ensuring that improvements are not only identified but systematically executed and tracked over time.

Production considerations

Multi-agent systems become more useful when designed thoughtfully.

Important considerations include:

  • Context-aware analysis: Agents should focus on areas relevant to the repository being analyzed.
  • Duplicate finding management: Overlapping findings should be consolidated to reduce noise.
  • Resource efficiency: Large repositories can increase token usage, execution cost, and runtime requirements.
  • Customization: Teams often need the ability to align audits with their own engineering standards and governance practices.
  • Agent scope: Each agent should have a clearly defined responsibility. Poorly scoped agents often produce overlapping or low-value findings.
  • Shared context: Agents need enough repository context to reason effectively without unnecessarily duplicating information across multiple context windows.
  • Finding structure: Findings are easier to review when agents return a common structure such as:
    • Severity,
    • Location,
    • Finding,
    • Evidence,
    • Impact, and
    • Recommendation.
  • Conflict resolution: Different agents may occasionally recommend conflicting actions. For example:
    • Performance optimization may conflict with maintainability goals.
    • Security recommendations may increase operational complexity.

    Human review is still required to evaluate trade-offs.

  • Confidence levels: AI-generated findings should be treated as recommendations supported by evidence rather than as guaranteed facts.

These design choices often have a significant impact on the usefulness of the final output.

Traditional code analysis vs. multi-agent AI code review

Traditional analysis Multi-agent AI audit
Rule-based checks AI-driven reasoning
Predefined conditions Domain-specific analysis
Deterministic results Context-aware recommendations
Usually focused on individual issues Can connect findings across domains
Lower computational cost Potentially higher execution cost

Multi-agent AI systems are not intended to replace static analysis, security scanners, tests, or human code review.

Instead, they complement these tools by adding broader repository-level reasoning and prioritization.

Where multi-agent code audits deliver value

Multi-agent audits are particularly useful for:

  • Large enterprise repositories,
  • Multi-service and microservice applications,
  • Teams adopting AI-assisted development,
  • Security-sensitive software, and
  • Organizations with multiple contributors and code owners.

As complexity grows, specialized analysis can become increasingly valuable.

Limitations worth understanding

Multi-agent AI systems are powerful, but they are not replacements for experienced engineers.

Important limitations include:

  • AI-generated findings can be incomplete or incorrect.
  • Multiple agents may produce conflicting recommendations.
  • Large repositories can increase execution costs.
  • Security-sensitive changes still require human validation.
  • Business and architectural decisions remain human responsibilities.
  • Multiple agents can increase token usage and execution costs.
  • Broader repository analysis may require additional context management and orchestration.
  • More agents do not automatically result in better outcomes.

The purpose of these systems is not to eliminate developers from the process. Their value lies in helping developers spend less time finding problems and more time solving them.

Frequently Asked Questions

What is a multi-agent AI system?

A multi-agent AI system uses multiple specialized agents that collaborate to analyze, monitor, or improve a shared environment.

How is it different from traditional code review tools?

Traditional tools often focus on specific checks in isolation.
Multi-agent systems combine multiple perspectives and help organize findings across different quality domains.

Can multi-agent audits replace human code reviews?

No. They work best as a complement to human reviews by identifying risks and opportunities earlier in the development process.

Are multi-agent AI systems useful for large repositories?

Yes. They can be particularly useful for large repositories, provided the system uses appropriate scoping, context management, orchestration, and resource controls.

What problems do multi-agent AI systems help address?

They can assist with technical debt, architectural drift, testing gaps, security risks, performance bottlenecks, and code quality management.

Take control of technical debt with Multi-agent AI systems

Generating code is becoming easier every day. Keeping code secure, maintainable, performant, and aligned with engineering standards remains a growing challenge.

Multi-agent AI systems help address that challenge by enabling specialized analysis across architecture, security, performance, testing, and code quality.

Rather than relying solely on periodic reviews, teams can apply structured repository analysis to identify risks earlier and prioritize improvements more effectively.

As AI-assisted development accelerates, successful teams will increasingly combine faster code generation with stronger code governance practices.

Syncfusion Code Studio allows teams to create specialized AI agents, run repository-wide audits, and turn findings into actionable improvement plans aligned with their engineering standards.

The goal isn’t autonomous code governance. It’s giving engineering teams deeper visibility into repository health and technical debt so they can make smarter decisions at scale.

Ready to see what your repository is really telling you?

Start your free trial or book a demo to explore AI-powered repository audits with Syncfusion Code Studio.

Read the whole story
alvinashcraft
27 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Add Observability to a Genkit AI Agent with Progress Agent Engineering

1 Share

Observability gives us a window into how an AI agent is working throughout its run, allowing us to improve its performance. Here’s how to add observability to a Genkit agent with Progress Agent Engineering.

Building an AI agent is easier today than it was a few years ago. We can connect a model, give it a prompt, add a few tools and get a useful answer quickly.

The difficult part starts after the first successful demo. What happened inside the agent? Why did it choose that answer? Which tool was slow? Did the prompt cause the problem? Is the model using too many tokens? Where is the request spending money?

As a frontend developer, I am used to opening DevTools and adding a console.log. I can check a network request and quickly see a 200 or a 500 response. That gives me a good starting point when a web application has a problem.

An AI agent is different. The request can be successful, but the answer can still be wrong, slow or too expensive. A 200 response only tells us that the request finished. It does not tell us what the model did, which tools it called or why it made that decision. At that point, our usual DevTools are not enough.

Adding a model is not only a model problem. There are several pieces to check:

  • The prompt may not give the model enough context.
  • A tool may return incomplete or incorrect data.
  • The model may call the wrong tool or call it more than once.
  • A second model call may change or misunderstand the tool results.
  • The full request may use more tokens and cost more than expected.

To understand these problems, we need to see more than the final answer and more than a network status code. We need visibility into the complete agent run. We need observability.

How We Will Explore Observability in a Run

The best way to learn is with a working example. Today, we’re going to build a Genkit travel agent. Genkit helps us build the agent, connect it to Gemini and give it tools.

Our example is a Family Travel Planner. It receives a destination, the number of adults and children, the children’s ages and a budget preference. It then creates a structured travel plan with flights, a hotel, activities and an estimated cost.

We will run the same agent before and after adding observability. This will give us a simple baseline and help us see what observability adds to the application. For our run, the agent will plan a family trip to Ibiza and use three tools for flights, hotels and activities. We’ll keep the logic unchanged to compare two runs fairly.

Without observabilityWith observability
We only see the final answerWe see the full agent run
Tool calls are hiddenEach tool call is visible
Token use is unknownInput and output tokens are shown
Cost is unknownCost is shown in the dashboard
Errors are hard to findWe can find the step that failed

By the end, we will have a native Genkit trace in an observability dashboard. The trace will show the generation, model calls, tool calls, token use and cost.

The Scenario

Imagine that you join a company that has started adding AI to its product. The team has a travel agent that creates plans for families. It uses a model, a prompt, and tools for flights, hotels and activities.

At first, everything works well. Then users start reporting that the agent is giving worse answers. Someone says that the model is not working anymore. Another developer remembers that the prompt changed yesterday. The data used by one of the tools was also updated this week.

Now we have several possible causes:

  • The model may be having a problem.
  • Someone may have changed the prompt.
  • A tool may be returning different or incomplete data.
  • The agent may be using more model calls than before.

How can we know where the problem is? Looking only at the final answer is not enough. We need to see the complete request and compare each step.

Why Do We Need Observability?

Console logs can tell us that a request started or failed. They usually do not show the complete relationship between the model and its tools. Observability gives us that context in one trace.

With observability, we can:

  • Find the exact model or tool span that is slow.
  • Inspect tool inputs and outputs when the final answer is wrong.
  • See how many input and output tokens a request uses.
  • Track the cost of individual requests and the application over time.
  • Find failed steps without reproducing the complete request locally.
  • Compare traces after changing a prompt, model or tool.

What problem does this solve for us as developers? It turns “the agent gave a bad answer” into a more useful question, such as “Did the hotel tool return incomplete data?” or “Did the second model call receive the tool results?”

When Should We Add It?

Observability is especially useful when an AI feature has more than one moving part. Consider adding it when:

  • An agent calls tools or other services.
  • A request includes multiple model calls or steps.
  • Response time, token usage or cost matters.
  • You need to investigate errors in development or production.
  • You are changing prompts or models and want to compare results.

For a very small experiment that makes one model call, a console log may be enough. For a real application, observability should be added early. It gives us a baseline before traffic grows and before debugging becomes harder.

The travel planner is a good example because one request combines three tools and one or more model calls. That makes the difference between a final answer and an inspectable execution path easy to see.

Why Not Use Only the Model Provider’s Tools?

Genkit and model providers such as OpenAI can provide useful logs and monitoring features. These tools are a good place to start, especially when we are testing one provider and one simple model call.

Things get harder when the application grows. An agent can use a framework, several tools and multiple model providers. If each provider gives us a different view, we have to jump between dashboards to understand a single request. We may see the model call in one place and the tool call somewhere else.

This can also tie our observability to one provider. If we change the model, add a new provider or move from Genkit to another framework, we may need to change our monitoring setup too.

That is why an observability platform that is independent from the model provider and the application framework can help. The goal is to keep one view of the complete request, even when the technology behind the request changes.

What Is Progress Progress Agent Engineering?

Progress Agent Engineering is an observability platform that helps us see what happens inside an application while it runs. For an AI application, it collects information about model calls, tool calls, errors, duration, token usage and cost. It then presents that information in a dashboard where we can inspect a request from beginning to end.

Think of it as a flight recorder for an AI agent. The final travel plan is the result, but the trace shows the complete journey: which tools were called, which model calls were made, how long each step took and what data each step produced.

Progress Agent Engineering can be used with applications built in TypeScript/JavaScript, Python and .NET. The idea is the same in each case: add the SDK to the application, collect the important events and inspect them in the dashboard.

That is the approach we will use here. The agent runs with Genkit and Gemini, but Progress Agent Engineering gives us one place to inspect the full execution. We use the Progress TypeScript SDK because this example is a Node.js application. The same idea can be applied with the Python or .NET SDKs in applications written in those languages.

A Few Observability Terms

Before we look at the features, let’s make the vocabulary simple.

  • Think of a trace as the complete receipt for one agent request.
  • A span is one line on that receipt, such as a model call or a tool call.
  • Instrumentation is the code that watches those operations and creates spans.
  • Telemetry is the information produced by that instrumentation and sent to Progress.
  • Content tracing means that prompts, responses and tool data can also be recorded. It is useful for debugging, but it can contain private information, so use it carefully.

What problem does observability solve for us as developers? Instead of guessing why a request was slow, expensive or incorrect, we can inspect each span that produced the result.

Progress Agent Engineering Features Used in This Project

Progress Agent Engineering includes several observability features. In this tutorial, we use these features:

  • Tracing: Groups everything that happens during one request into a single trace. This includes the Genkit flow, model calls and travel tools.
  • Native Genkit span processing: Understands the spans created by Genkit and connects them to the trace. We do not need to add a new span for every operation.
  • Trace details: Lets us open a trace and inspect its individual spans, status, duration, model, provider and tool data.
  • Token and cost tracking: Shows input tokens, output tokens, total tokens and the estimated cost of the model calls.
  • Agents page: Provides an application-level view with the active agent, span count and accumulated cost.
  • Content tracing: Can record prompts, model responses and tool data so we can investigate an incorrect result. This data may be private, so enable it only when it is safe to do so.

The SDK also provides the start and stop methods used by the application. Observability.instrument() starts monitoring before Genkit loads, and Observability.shutdown() sends the remaining telemetry before this short command exits.

These features work together. Monitoring collects the events; tracing groups them. The dashboard helps us inspect the events, and usage data helps us understand the running cost of the agent.

The main files are:

src/agent.ts                 Genkit agent, tools, and output schema
src/app.ts                   Command line application
src/tools/                   Mock flight, hotel, and activity data

We will not change the agent logic. Both runs must use the same code.

How the Pieces Fit Together

The project repository contains the complete runnable agent; we only need to change the application startup and shutdown flow:

  • src/agent.ts defines the Genkit agent, its output schema, and its tools.
  • src/tools/ contains the mock travel data.
  • src/app.ts reads the command-line arguments, calls the agent, and prints the result.
  • bootstrap.ts starts Progress, loads the application, and flushes telemetry before the process exits.

The full agent and tool code is already part of the project repository. We don’t repeat it here because observability shouldn’t require rewriting the travel logic.

Get the Project

First, clone the project repository and move into its folder. Replace <repository-url> with the URL of the repository that contains this example:

git clone https://github.com/danywalls/genkit-progress-observability.git
cd genkit-progress-observability
git branch --all

You should see the article-start and main branches. The first branch contains the agent before observability.

Before You Start

You need a Google AI Studio API key for both branches. The main branch also needs a Progress Agent Engineering integration key. Create a local .env file in the project root:

GOOGLE_API_KEY=your-google-api-key
OBSERVABILITY_API_KEY=ac_p_your-integration-key
OBSERVABILITY_APP_NAME=family-travel-planner
GEMINI_MODEL=gemini-3.5-flash-lite

The article-start branch uses only the Google key to call Gemini. The main branch uses both keys. The .env file is ignored by Git, so it will not be committed.

Run the Agent Without Observability

Start from the article-start branch:

git checkout article-start
npm install
npm start -- Ibiza 2 2 moderate 7 5

The terminal shows the travel plan and the tool messages. This is useful, but it does not answer important questions:

  • Did the model call all three tools?
  • Which model call took most of the time?
  • How many tokens did the request use?
  • How much did this request cost?
  • Which step should we inspect when the answer is wrong?

The command prints a travel plan like this:

Ibiza Family Travel Planner - Genkit
Planning a family trip to Ibiza...
[tool:find-flights] Searching flights to Ibiza for 4 passengers
[tool:find-hotels] Searching hotels in Ibiza
[tool:find-activities] Finding activities in Ibiza for 2 adults and 2 kids
[agent:stream] model started producing a structured travel plan
TOTAL ESTIMATED COST: €1,350

The agent works, but it is still a black box. We need more information about the run. Now let’s add the observability bootstrap around the same application.

Add Progress Agent Engineering Observability

Add Progress Observability

If you have just run the agent from article-start, you already have the part that matters most: a working Genkit flow. We are not going to redesign that flow or add tracing calls around every tool. We are going to place a small observability layer around the application that is already there.

Here is the change we are about to make:

  1. Move the command entry point to bootstrap.ts.
  2. Start Progress before src/app.ts loads Genkit and the Google AI plugin.
  3. Let the existing agent run as it does today.
  4. Shut Progress down after the request so the last spans are sent to the dashboard.

Genkit already creates spans for the flow, model calls and tools, and Progress 3.1.1 knows how to process those Genkit spans. Its native Genkit support also maps tool names, arguments and results, as well as model token usage, into the trace. Our job is to put the start and end of the observability lifecycle in the right places. Once that is clear, the integration is easier to follow: one file controls startup, the existing application runs in the middle, and one shutdown call closes the run.

1. Install the Integration

Switch to the branch that contains the observability setup and install its dependencies:

git checkout main
npm install

The main branch changes the entry point in package.json from tsx src/app.ts to tsx bootstrap.ts. This makes bootstrap.ts the first file executed by the command. That order matters: if src/app.ts loads Genkit before Progress starts, some libraries may already be initialized and their operations may not be captured.

2. Start Progress Before Loading the Application

Create bootstrap.ts in the project root. Its job is to prepare the environment, validate the required keys, start the SDK and only then load the application:

import '@progress/observability/register/hooks';
import 'dotenv/config';

import { Observability, ObservabilityInstruments } from '@progress/observability';

const apiKey = process.env.OBSERVABILITY_API_KEY;
if (!apiKey) throw new Error('OBSERVABILITY_API_KEY is not set');

if (!process.env.GOOGLE_API_KEY) {
  throw new Error('GOOGLE_API_KEY is not set');
}

await Observability.instrument({
  appName: process.env.OBSERVABILITY_APP_NAME ?? 'family-travel-planner',
  apiKey,
  instruments: new Set([ObservabilityInstruments.GOOGLE_GENERATIVEAI]),
  traceContent: true,
});

await import('./src/app.js');

Let’s follow the file in the order Node executes it.

First, the hooks are registered:

import '@progress/observability/register/hooks';

This gives Progress a way to observe supported libraries as they are loaded. It needs to appear before the application imports Genkit or the Google AI plugin. If those libraries are loaded first, their initialization may happen before Progress has installed its instrumentation.

Next, we load the values from .env:

import 'dotenv/config';

After this import, process.env contains GOOGLE_API_KEY, OBSERVABILITY_API_KEY and OBSERVABILITY_APP_NAME. We read the Progress key and check both required keys before starting the application. Failing here gives us a clear configuration error instead of a request that cannot be traced or sent to Gemini.

Now we start the observability SDK:

await Observability.instrument({
  appName: process.env.OBSERVABILITY_APP_NAME ?? 'family-travel-planner',
  apiKey,
  instruments: new Set([ObservabilityInstruments.GOOGLE_GENERATIVEAI]),
  traceContent: true,
});

Think of instrument() as the point where we turn observability on for this process. It prepares the telemetry pipeline and configures what Progress should capture:

  • appName groups this application’s traces under family-travel-planner in the dashboard.
  • apiKey identifies the Progress integration that receives the telemetry.
  • instruments enables the Google Generative AI instrumentation used by the Genkit Google AI plugin.
  • traceContent includes prompts, responses and tool data in the trace. That helps us debug an incorrect result, but the captured content may contain sensitive information.

The call is asynchronous because the SDK needs to prepare that pipeline. await makes the bootstrap wait until the setup has completed. We do not want the agent to start while the observability layer is still initializing.

Only after instrument() finishes do we load the application:

await import('./src/app.js');

This is a dynamic import, rather than a static import at the top of the file, for one reason: it keeps Genkit and the agent from loading too early. From this point on, src/app.ts runs the same parseArgs(), planTrip() and output code as before, but its Genkit flow and model calls are now observed by Progress.

At this point, Progress can observe the Genkit flow, but a short-lived CLI process can exit before its telemetry is transmitted. We therefore need an explicit shutdown step.

3. Close the SDK After the Agent Run

Open src/app.ts. Keep the argument parsing, planTrip() call and output formatting from the first run. Add the Observability import and call shutdown() from the existing finally block:

 import { planTrip } from './agent.js';
+import { Observability } from '@progress/observability';
 
 // The existing main() function still parses arguments, calls planTrip()
 // and prints the travel plan.
 try {
   // Existing planTrip() call and output.
 } catch (error) {
   console.error('Error generating travel plan:', error);
   process.exitCode = 1;
 } finally {
+  await Observability.shutdown();
 }

finally runs after both a successful and a failed request. That makes it the correct place to flush the remaining spans. The agent logic remains unchanged: it still calls the same tools and produces the same structured travel plan. The only new behavior is that the SDK closes after the run has produced its result.

The resulting application structure is:

bootstrap.ts
  1. Load hooks and environment variables
  2. Start Progress
  3. Load src/app.ts
       4. Run the existing Genkit agent
       5. Shut down Progress in finally

This separation is useful because bootstrap.ts controls initialization while src/app.ts controls one agent run. It also explains why the application does not need custom tracing code around find-flights, find-hotels or find-activities.

At the end of the integration, the final version is organized like this:

bootstrap.ts
  - Loads the Progress hooks and environment variables
  - Validates the API keys
  - Starts Observability.instrument()
  - Loads src/app.ts only after instrumentation is ready

src/app.ts
  - Parses the command-line arguments
  - Calls the existing planTrip() function
  - Prints the travel plan
  - Calls Observability.shutdown() in finally

The important change is the execution order. Progress starts before the agent, the existing Genkit code runs without custom tracing changes, and the SDK shuts down after the request. That is the complete integration we need for this command-line application.

The application now starts Progress before loading Genkit and shuts it down after the request finishes. The agent and its tools are still unchanged. Run the same command as before:

npm start -- Ibiza 2 2 moderate 7 5

The terminal output is similar to the first run. The difference is that Progress now receives the full trace.

The command prints the travel plan and confirms that telemetry was sent:

Ibiza Family Travel Planner - Genkit + Progress Observability
Planning a family trip to Ibiza...
[agent:stream] model started producing a structured travel plan
TOTAL ESTIMATED COST: €1,350
Check your traces at: https://observability.progress.com
[Observability] Observability shutdown completed successfully

The exact travel plan and total may be different in your run. Model responses are not always identical, even when we send the same input. For this comparison, focus on the execution path and the trace data, not on matching every word or number.

Let’s read the trace in the Progress observability platform and go to Observe > Tracing. Find the family-travel-planner service.

Open the latest trace. A normal trace looks like this:

The exact number of spans can change when the model makes a different number of calls. The important point is that the trace shows the whole run in one place.

When you open the trace, read it in this order:

  1. Check the trace status. Did the request finish successfully?
  2. Check the total duration. Is the request slower than expected?
  3. Open the tool spans. Did the agent call the expected tools, and did they return useful data? In the Tool Details panel, Progress shows the tool name, the JSON arguments sent by Genkit and the result returned by the tool.
  4. Open the model spans. Which model was used, and how long did each call take?
  5. Check tokens and cost. Is the request using more resources than expected?

To answer these questions, check model usage. The trace summary shows input tokens, output tokens, total tokens, input and output cost, total cost and request duration.

In one real test run, the trace showed 7 spans, 2,411 total tokens and a total cost of $0.0021. Token counts can vary between runs. The model span also shows the model name and provider information. The generate span is a Genkit flow span, so it may not have model provider data.

Finally, the Agents page gives us a wider view of the application. It shows the active agent, span count and accumulated cost.

At this point, we have followed one request from the command line to its trace. The trace connects the Genkit flow, model calls and tools, while the summary shows usage and cost and turn that data into practical debugging decisions.

Why This Helps

The second run gives us information that the first run did not provide.

If the answer is slow, we can check the duration of the model and tool spans. If the answer is wrong, we can inspect the tool output and the model input. If the cost grows, we can compare token use between runs.

This is the main value of observability. It changes a guess into information we can use, the best part is the. The project does not add custom code to create model spans or map Genkit token fields, tool names, arguments or results because Progress Observability 3.1.1 handles the Genkit data directly.

This keeps the application simple:

  • Genkit runs the agent and its tools.
  • Progress records the Genkit trace.
  • The dashboard shows model use and cost.
  • The application code stays focused on travel planning.

Conclusion

The agent worked before observability. That was enough to produce an answer, but not enough to understand it or identify the next improvement.

After adding observability, we can see the complete run: the generation steps, tools, model calls, tokens and cost. This gives us a clear place to look when the agent is slow, when the answer is wrong or when the cost is growing. We can use that information to improve the prompt, fix a tool, change the model or remove unnecessary model calls.

The integration is also small. We start the SDK before Genkit loads, run the existing agent, and shut the SDK down when the request finishes. We do not need to rewrite the agent logic or add custom code around every tool call. That makes it practical to add observability to an agent that is already running in a product.

Remember that the observability layer is not tied to one model provider. We used Genkit and Gemini, but the same idea can remain in place if the application adds another provider or changes its model. Progress gives the team one view of the request while the technology behind that request evolves.

Happy observability!!

Read the whole story
alvinashcraft
33 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Understanding Device Bound Session Credentials (DBSC)

1 Share

In this post I provide an introduction to Device Bound Session Credentials (DBSC). I describe the problem they're trying to solve, how the protocol works, and what you need to do to to support them.

What problems are Device Bound Session Credentials (DBSC) trying to solve?

Web applications need a way to authenticate. One of the most common (and reasonable) approaches is to rely on long-lived cookies for user authentication. With modern enhancements like same-site cookies, many of the historical vulnerabilities associated with cookie authentication are now less of an issue.

However, a fundamental issue with authentication session cookies remains: these are bearer tokens. That is, there's no way to prove you "own" the cookie; if anyone else has the cookie, there's nothing to stop them using it.

When you hear about "bearer" tokens, you might typically think about JWT tokens which are commonly sent in headers. But all "bearer" means is "anyone in possession of the token can use it", which applies to cookies too.

This leads to a risk of "cookie theft" and to "session hijacking", where if an attacker finds a way to read your authentication cookie, they're free to use it for malicious behaviour, on any computer. They don't have to find a way to send requests from the victim's computer; as long as they can extract the cookie, they can use it from their own machines without issue.

Device Bound Session Credentials (DBSC) aim to provide a defence against this weakness, by allowing a server to verify that the authentication cookie is being used on the same machine it was issued to. This makes session hijacking harder, as you can no longer simply extract the cookie and use it elsewhere; such a cookie would not be valid.

This verification works by having the browser sign requests using a private key that is stored in a Trusted Platform Module (TPM). The key is never exposed outside of the TPM, so you can be sure that if a request is signed with the same key, it's coming from the same machine.

DBSC is also meant to provide progressive enhancement, and doesn't require changing how web applications do their authentication. A server can request that browsers use DBSC where possible, and if the browser supports it, it will seamlessly opt in. A browser that doesn't support DBSC won't recognise the opt-in header, and so will continue as it did before; there's no breaking changes.

How does DBSC work?

In this section, I describe the overall flow for a server and browser that implement DBSC. This is a very general overview, and there are no doubt lots of subtleties in any actual implementation (e.g. see Scott Helme's post about edge cases and bugs he ran into). The goal is to follow the various additional requests happening, so that you can better understand what's happening if you choose to implement DBSC in your applications.

Implementing DBSC

Implementing DBSC in a server should be relatively simple (when compared to some other security measures). In brief, an application needs to:

  • Add a Secure-Session-Registration header to the response when a user logs in. This advertises to the browser that the server supports DBSC.
  • Implement a "session registration" endpoint that registers the browser session and switches to short-lived cookies.
  • Implement a "refresh" endpoint that validates the browser still has access to the key, and refreshes the short-lived session cookie.

An important aspect of DBSC is that you don't need to change how you do authentication and validation in general. You need to add the endpoints above, and handle the implications, but overall your authentication code should just keep working as it currently does.

Also, this flow is optional. If a browser doesn't recognise the Secure-Session-Registration header, then it doesn't support DBSC, and so nothing happens. Your application continues to work just as it did before you implemented DBSC.

Now that you know you only have to implement a couple of endpoints, let's look at the end-to-end flow of DBSC.

The DBSC flow

The overall flow of requests, after a user logs in to a server that supports DBSC is shown in the following diagram:

The DBSC flow

Note that there are alternative integration patterns available, so you may see other variations of the above.

1. The login response includes a Secure-Session-Registration header

The first step is for the server to advertise to the browser that it supports DBSC. When a user logs in, as well as the usual authentication cookie, the server adds an additional header to the response, Secure-Session-Registration:

HTTP/1.1 200 OK
Secure-Session-Registration: (ES256 RS256); path="/dbsc/registration";challenge="4a96e2cab06e76e2366b5e802bfcbabfe81e52c81abfcc71afc97010157ae9bd"
Set-Cookie: MyAuthCookie=CfDJ8Apn9==; path=/; samesite=lax; httponly

The structure of the that Secure-Session-Registration cookie is as follows:

  • (ES256 RS256): These are the supported algorithms that can be used for signing by the browser. In this case, the supported algorithms are ECDSA P-256 and RSASSA-PKCS1-v1_5.
  • path="/.well-known/dbsc/registration": This is the "registration" path that the browser should use to register new DBSC credentials.
  • challenge="<somevalue>": A random value that the browser will sign and send as part of the registration call

There are some additional fields that could be included in the Secure-Session-Registration header, but which are not required by all implementations, such as authorization or provider_key, but I'll ignore those here for now. You can read more about them in the specification here.

When the browser receives the response, and it sees the Secure-Session-Registration header, this is the trigger that it should register a key with the DBSC endpoint your application exposes.

2. The browser sends a request to the registration endpoint

This is the crucial step in the flow. On receiving the Secure-Session-Registration header, the browser does the following:

  • Generate a new public-private key pair in the TPM/secure enclave
  • Sign the provided challenge with the private key.
  • POST a request to the specified path, containing the signed challenge and the public key, as a JWT.

So the browser sends something like this:

POST /dbsc/registration HTTP/1.1
Cookie: MyAuthCookie=CfDJ8Apn9==
Secure-Session-Response: eyJhbGciOiJFUzI1NiIsImp3ayI6eyJjcnYiOiJQLTI1NiIsImt0eSI6IkVDIiwieCI6ImxITjNhci13bFZTU0FkeThPSlhxeGhId0JpdXVyMnJUbG1ieGNnaW05X28iLCJ5IjoiemZfeTc5cDhycGI3enRuUUdaMjV4UFhfTHFfcjdvMUNoOWp2bmU3MHRKNCJ9LCJ0eXAiOiJkYnNjK2p3dCJ9.eyJqdGkiOiI0YTk2ZTJjYWIwNmU3NmUyMzY2YjVlODAyYmZjYmFiZmU4MWU1MmM4MWFiZmNjNzFhZmM5NzAxMDE1N2FlOWJkIn0.MO28tc5E-OwA7IlKTAMe1yCGkRA_b8ljLVpP0gnc-jky7g1dcw2CrYREB0KWoE_ae5ixCjJZc4IEcVwfnP_qKA
Content-Length: 0

As you can see, this is posting to the Path provided in the original header, and it includes the original authentication cookie. The Secure-Session-Response contains a base64 encoded JWT. If you plug that into a decoder, you'll see that it contains:

  • A header, indicating the algorithm used for signing, along with the public key
  • The payload, which just contains jti, followed by the challenge sent in the header
  • The signature, which is generated from the header and payload, and the private key

The decoded header above looks like this:

{
  "alg": "ES256",
  "jwk": {
    "crv": "P-256",
    "kty": "EC",
    "x": "lHN3ar-wlVSSAdy8OJXqxhHwBiuur2rTlmbxcgim9_o",
    "y": "zf_y79p8rpb7ztnQGZ25xPX_Lq_r7o1Ch9jvne70tJ4"
  },
  "typ": "dbsc+jwt"
}

while the body looks like this:

{
  "jti": "4a96e2cab06e76e2366b5e802bfcbabfe81e52c81abfcc71afc97010157ae9bd"
}

It's then up to the server to handle this request.

3. The server validates the request, and returns short-lived credentials

When the server receives this request it must first validate the JWT in the Secure-Session-Response header, and confirm that

  1. The JWT signature is correct and valid.
  2. There is a valid authentication cookie for the user.
  3. The challenge in the JWT matches the one included in the original Secure-Session-Registration header.

If all those are true, then the server should do the following:

  1. Generate a new DBSC session ID, and associate it with the current user.
  2. Replace the existing authentication cookie with a short-lived cookie.
  3. Return instructions in the response for how to refresh the short-lived cookie, and where the cookie should be used.

If all goes well, the response will look something like this:

HTTP/1.1 200 OK
Content-Type: application/json
Set-Cookie: MyDbscCookie=abc123; Path=/; Secure; HttpOnly; SameSite=Lax; Max-Age=300

{
  "session_identifier": "199d6f60681",
  "refresh_url": "/dbsc/refresh",
  "scope": {
    "origin": "https://example.com",
    "include_site": false
  },
  "credentials": [
    {
      "type": "cookie",
      "name": "MyDbscCookie",
      "attributes": "Path=/; Secure; HttpOnly; SameSite=Lax"
    }
  ]
}

Let's dig through each part of this response:

  • The long-lived MyAuthCookie cookie is no longer present
  • A new MyDbscCookie cookie is now set, with a short expiration time (5 minutes). This now serves as the authentication cookie for the user.
  • The body contains a session_identifier, which is used to associate this DBSC session with the current user. The browser will include this whenever it needs to refresh the short-lived auth cookie.
  • The body contains the refresh_url, which is the path the browser must hit to obtain a new short-lived auth cookie.
  • The scope says where the short-lived cookie is valid, and where it should not be used
  • Finally, the credentials section provides details about the exact cookie that this config applies to.

Note that the only valid value for "type" is cookie, so this is likely just future-proofing at work.

After receiving the response, the browser will then use these credentials for all subsequent requests. The server must use the short-lived cookie in place of the "normal" authentication cookie, and everything works as "normal" aside from that.

Of course, very soon, that cookie is going to expire, so it's important for the browser to be able to refresh these credentials.

4. The browser refreshes the credentials in the background

When the browser needs to send a request to your application, and the short-lived cookie has expired, the browser needs to get a fresh instance of the cookie. In general, browsers will likely try to refresh before the cookie expires but if they don't, then they will defer a user request until after the refresh request.

The browser first sends a request to the refresh endpoint, including the DBSC session ID:

POST /dbsc/refresh HTTP/1.1
Sec-Secure-Session-Id: 199d6f60681

The server then generates a new challenge for the browser to sign in the Secure-Session-Challenge header and returns a 403 response:

HTTP/1.1 403 Forbidden
Secure-Session-Challenge: "4524d32ab2b9";id="199d6f60681"

Next, the browser creates another signed JWT, containing the the challenge value with the same private key as it used to register the session originally. It then sends another request to the same refresh endpoint, but this time with the JWT included in the Secure-Session-Response header:

POST /dbsc/refresh HTTP/1.1
Sec-Secure-Session-Id: 199d6f60681
Secure-Session-Response: eyJhbGciOiJFUzI1NiIsImp3ayI6eyJjcnYiOiJ...

The server then has to make sure everything matches up:

  1. The JWT signature is correct and valid.
  2. The session ID is a known current session
  3. The challenge contained in the JWT matches the challenge sent to the browser

If all those checks pass, the server creates a new short-lived cookie, and responds with a 200 OK:

HTTP/1.1 200 OK
Set-Cookie: MyDbscCookie=def456; Path=/; Secure; HttpOnly; SameSite=Lax; Max-Age=300

When the browser receives the request, it replaces the expired short-lived cookie with the new one, and continues with the deferred request.

And that's the complete DBSC covered. As the cookie expires, the browser keeps calling the refresh endpoint to create new short-lived cookies.

Can you use DBSC? Is it worth it?

So, as an application creator, should you support DBSC? In general, I think the answer is "probably", because it doesn't really have an obvious downside, and it protects your users from session hijacking.

Right now, only Chromium implements DBSC, so while that has a big reach, it's far from ubiquitous. The good news is that you should be able to implement DBSC, and then it will only kick in if the user's browser supports it. If the browser doesn't support DBSC, then it will ignore the Secure-Session-Registration header entirely, and the browser simply uses the normal long-lived/session cookie to authenticate with the browser as normal.

But if the browser does support DBSC, you get all the benefits that brings. Namely, protection against session-hijacking, by converting to short-lived credentials instead. And it shouldn't require sweeping changes to your app. So why not?

There are some potential difficulties with the implementation, as described in this post from Scott Helme, which can be very problematic in some cases. For that reason, I'd suggest waiting for the framework to implement it. On a more minor note, I found that my ad blocker completely blocked the DBSC flow 😅

So in conclusion: yes, implement it for extra security, but maybe wait for a canonical implementation in your language/framework of choice first!

Resources

I wrote this post based on reading a bunch of others, I recommend the following to get a better understanding of the feature:

Read the whole story
alvinashcraft
41 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Tell agents the why, not just the how

1 Share

Early AI agents were basically enthusiastic idiots. Working with them required you to tell them precisely what you wanted them to do (for instance, “method A exists on class B, please add an equivalent method to classes C through F”). Otherwise they’d go off and do entirely the wrong thing. But as AI agents have improved, this has changed.

When frontier models go off and do the wrong thing today, they don’t do it because they’re confused, they do it because they make an incorrect assumption about your goals or priorities. For instance, when GPT-6-Astra thinks it’s writing code for itself, it will produce minified code. It’s perfectly capable of writing human-readable code — at least in Golang, where I’ve produced several thousand lines of acceptable code with the model — but you have to tell it that humans will be reading the code1.

This is the main piece of advice I want to give most people I see prompting agents: give the agent context on your priorities, not just on the specific task you want them to do. Here’s a prompt I recently used as the starting point for Deckard.

Hello. You should have Runpod access via MCP (if not, tell me and I’ll fix it).

I have the long-term goal of building a local program or browser extension to automatically scan pages I load for AI content and hide it. I have the short-term goal of figuring out the best AI detection model I can run on my macbook without killing my battery or making it hot, and (relatedly) figuring out how to run the model most efficiently. My guess is that Pangram’s EditLens 3B or the smaller Roberta model might be a good place to start, though they might require quantizing and will definitely require some work to make them run as efficiently as possible on my macbook.

I would like you to use my Runpod account to start answering these questions. Eventually we’ll move to doing things on this macbook pro, but my hope is that Runpod can help with some experiments that are too hot/long/slow to run locally. You are a smart model; if you can see a better way to achieve my goals, please let me know and we’ll talk about it. Good luck.

About half of this prompt is sharing broad context, such as the overall project I’m aiming for, the fact that it’s for me personally and not for work, and my priorities (e.g. keeping the laptop cold). If I had written an explicit spec, I would have missed a bunch of improvements: for instance, using native messaging for the local model, or choosing the Gradient model instead of EditLens.

I do the same thing for work, but typically with a stronger emphasis on my technical values. I often write a paragraph spiel explaining the relative priorities of avoiding bugs, observability, fitting elegantly into the current code, performance, and so on. Note that I said “relative” priorities: I don’t simply list all of these things and say they’re important, I explicitly tell the model which of those I care less about and can therefore be traded off to better achieve the others.

Models are now smart enough to have meaningful input on your broader goals. If you’re just prompting them with a concrete technical spec, you are committing the same mistake as in the XY problem: asking expert advice without giving the expert the context it needs.


  1. Incidentally, you don’t have to tell it to write human code if it’s working in a human-authored codebase. It’s smart enough to pick up the style of the surrounding code.

Read the whole story
alvinashcraft
51 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

The Mediator pattern with MassTransit

1 Share

For years, when a .NET developer said "I need the mediator pattern," the answer was almost automatically MediatR. Jimmy Bogard's library became the default so quickly that a lot of teams never really questioned it, they just added the NuGet package and moved on.

That changed in 2025. MediatR moved to a commercial license, and suddenly "just add MediatR" wasn't a free decision anymore. That's a good moment to take a step back and look at what the mediator pattern actually is, and to point out something a lot of developers overlook: if you're already using MassTransit in your project, you already have a mediator implementation available. No extra package, no extra license to think about.

Remark: Be aware that MassTransit switched to a commercial license as well starting from version 9.

What is the mediator pattern?

The mediator pattern is about decoupling the sender of a request from the thing that handles it. Instead of a controller or service directly calling into a handler class, it sends a message to a mediator, and the mediator figures out which handler should process it.

The benefit: your calling code doesn't need to know about the handler's dependencies, its constructor, or even which assembly it lives in. You send a command or a request, and the mediator dispatches it. This is especially useful in a CQRS-style setup, where you want a thin API layer that just sends commands and queries without wiring up every handler by hand.

MediatR popularized this for in-process messaging in .NET. But the pattern itself doesn't require MediatR, it's just an interface (Send, Publish) and some dispatch logic behind it. And that's exactly what MassTransit's mediator gives you.

Two ways to get there

If you're starting from scratch, adding MediatR is still the most well-known route. But if MassTransit is already part of your stack (and for a lot of message-driven .NET systems, it is), you're pulling in a second in-process messaging abstraction that does almost the same thing your existing library can already do.

MassTransit ships with an in-memory mediator implementation, built on the same consumer model you already use for your bus. No transport, no broker, everything runs in-process, but you get the same IConsumer<T> shape you're used to from regular MassTransit consumers.

Remark: this is a meaningful mental shift if you're coming from MediatR. MediatR is handler-based: you implement IRequestHandler<TRequest, TResponse>. MassTransit's mediator is consumer-based: you implement IConsumer<T>, the same interface you'd use for a real bus consumer. That consistency is actually the whole point, your in-process and out-of-process messaging code looks the same.

Setting it up

Add the mediator in your service registration:

builder.Services.AddMediator(cfg =>
{
    cfg.AddConsumer<SubmitOrderConsumer>();
    cfg.AddConsumer<GetOrderStatusConsumer>();
});

Define your message contract like you would for any MassTransit message:

public record SubmitOrder
{
    public Guid OrderId { get; init; }
    public Guid CustomerId { get; init; }
}

And a consumer to handle it:

public class SubmitOrderConsumer : IConsumer<SubmitOrder>
{
    public async Task Consume(ConsumeContext<SubmitOrder> context)
    {
        var order = context.Message;

        // handle the command
    }
}

To send the command from a controller or minimal API endpoint, inject IMediator:

app.MapPost("/orders", async (SubmitOrder command, IMediator mediator) =>
{
    await mediator.Send(command);
    return Results.Accepted();
});

Request/response

Fire-and-forget commands are only half the story. If you need a response, MassTransit's mediator supports the request/response pattern too, using a request client created from IMediator:

public class GetOrderStatusConsumer : IConsumer<GetOrderStatus>
{
    public async Task Consume(ConsumeContext<GetOrderStatus> context)
    {
        await context.RespondAsync(new OrderStatusResult
        {
            OrderId = context.Message.OrderId,
            Status = "Processing"
        });
    }
}
var client = mediator.CreateRequestClient<GetOrderStatus>();
var response = await client.GetResponse<OrderStatusResult>(new GetOrderStatus { OrderId = orderId });

Same request/response shape you'd use with a real bus, just resolved in-process.

Remark: MassTransit dispatches asynchronously under the hood, but Send only completes once the consumer's Consume method completes. If the consumer throws, the exception propagates back to the caller, so your normal error handling around Send/GetResponse still works the way you'd expect.

Scoped vs singleton

IMediator is registered as a singleton, and by default each consumer gets its own DI scope. If you need several consumers in a pipeline to share the same scope (think: sharing a DbContext across steps), inject IScopedMediator instead. No extra configuration needed, it's available as soon as you've called AddMediator().

Do you lose anything?

Middleware like UseMessageRetry and UseInMemoryOutbox works with the mediator too, so you're not giving up MassTransit's pipeline behaviors just because you're not going over a transport. What you don't get: routing slip activities aren't supported through the mediator, but consumers, handlers, and sagas (including saga state machines) are fully supported.

If your project doesn't use MassTransit for anything else, MediatR (or a small hand-rolled mediator, plenty of examples of that around too) is still a reasonable choice. But if MassTransit is already in your dependency tree, you're one AddMediator() call away from not needing a second library for the same job.

That's it!

More information

Read the whole story
alvinashcraft
56 seconds ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories