Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162321 stories
·
33 followers

1.0.92

1 Share

2026-10-05

  • Add copilot config subcommands to list, read, set, and remove settings.
  • Add a pre-conversation Ctrl+E environment picker to switch between local and cloud runs
  • Entra-protected MCP servers can silently renew access-token-only credentials.
  • Legacy HTTP+SSE MCP connections no longer hang indefinitely when a message POST is never acknowledged; the acknowledgement is bounded by the server's configured timeout
  • Voice runtime install errors name why the download from nuget.org failed, not only the fallback feed's 401
  • Compaction keeps your latest prompt when requests exceed context limits
  • Large Anthropic requests rejected by provider size limits now retry after downscaling images or removing attachments
  • Custom agent model entries keep model-bound reasoning effort only when that model is selected
  • Shell tool calls now stream live stdout and stderr output reliably in the timeline
  • Usage reporting preserves provider-reported reasoning token totals when available
  • Sessions no longer slow to a crawl for minutes after the agent writes a very large file in one step
  • Sandboxed shells now withhold ambient GITHUB_TOKEN unless explicitly configured.
  • Plan usage reflects the current billing period after quota resets
  • Custom agents launched through ACP task calls now resolve and run correctly
  • Remote session resume now uses your configured GitHub auth for --resume and --connect
  • Entra sign-in falls back to browser auth when no broker is available and shows a manual URL when auto-open fails
  • copilot sandbox ca commands now respect --config-dir (including with -C).
  • Search commands avoid sandbox bypass prompts when access is already granted
  • Sandboxed uv commands can write to the uv cache by default when dev tool access is enabled
  • pnpm commands run in sandboxed sessions without lock-file permission errors
  • Pressing n repeatedly in the Sessions tab reliably creates each new session
  • Reverse search updates results when command history finishes loading at startup
  • MCP tools recover within the same turn when server instructions change
  • Changing experimental mode with /experimental or /settings takes effect after restart even when launched with an opposing experimental flag
  • Keyboard, paste, and mouse input now stay ordered and responsive during rapid interaction.
  • Sandboxed shell commands offer a network bypass prompt whenever the proxy blocks a destination
  • Sandboxed scripts that run Git now authenticate with masked credentials and SSH remote rewrites
  • Sub-agents keep working after you replace your GitHub authentication credentials
  • Sandboxed commands on Windows write temporary files to the granted temp directory, so tools that rename a temp file into place work
  • Prompt-mode sessions fire a single sessionEnd hook after Stop-hook continuations complete
  • Reconnect to remote MCP servers after idle Streamable HTTP sessions expire
  • Messaging a running background agent now steers its active turn at the next processing opportunity.
  • Context rollovers keep your latest requests in the recovery context.
  • Hide the automatic sandbox CA setup prompt on Windows accounts that cannot self-elevate
  • Copilot no longer describes a subagent as configured when the current session cannot run it. Previously a session that could not use rubber-duck still reported it as set up in /subagents, and asking for it failed instead of being skipped cleanly.
  • MCP tools continue working after OAuth reauthentication when tool definitions are unchanged
  • Select which account to use after Microsoft Entra sign-in, and let /logout sign out those OAuth sessions.
  • Improve first-run startup by extracting the bundled CLI package in a child process
  • Improve startup responsiveness when connecting many MCP servers at once
  • Canvas actions can now return images to the model in invoke_canvas_action.
  • Remove retired models from the model picker and supported CLI selections
Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

# Ink comes to WinUI 3 ✍️

1 Share

Handwriting is one of those inputs that either feels instant or feels broken. There is no middle ground.

You asked for it in #7426, "Is there a plan/schedule for adding InkCanvas into WinUI desktop?", and again in #10817, when apps moving over from UWP found the control simply was not there.

It is there now. WinUI 3 has InkCanvas and InkToolbar, available from Windows App SDK 2.5.4-experimental. Between them you get a surface that collects strokes and a tool strip that drives it, so a note app stops being a drawing engine you have to write yourself.

Pen

Pen strokes in blue, red and green at increasing widths

Pencil

Pencil strokes showing graphite texture

Highlighter

Translucent highlighter bands in yellow, cyan and pink

Here is the whole thing working. One canvas, one toolbar, no custom drawing code:

Ink Notes: choosing a pen colour, drawing against the ruler stencil, writing, highlighting and sketching

A canvas is one element

InkCanvas collects strokes on its own:

<muxc:InkCanvas x:Name="Ink" />

🎨 The toolbar is the second element

InkToolbar attaches to a canvas through TargetInkCanvas:

<muxc:InkToolbar x:Name="Toolbar"
                 TargetInkCanvas="{x:Bind Ink}"
                 HorizontalAlignment="Left">
    <muxc:InkToolbarBallpointPenButton x:Name="PenButton" />
    <muxc:InkToolbarPencilButton />
    <muxc:InkToolbarHighlighterButton />
    <muxc:InkToolbarEraserButton />
    <muxc:InkToolbarStencilButton x:Name="StencilButton" />
</muxc:InkToolbar>

💾 Save a picture that is still ink

GifWithEmbeddedIsf writes an ordinary GIF with the stroke data tucked inside it, so loading the file back returns editable strokes:

// GifWithEmbeddedIsf writes a normal GIF that any viewer can open, with the stroke data
// travelling inside it. Loading the file back returns editable strokes, not a flat picture.
private async Task<int> SaveAndReloadAsync(string path)
{
    var stream = new InMemoryRandomAccessStream();
    await Ink.InkPresenter.StrokeContainer.SaveAsync(
        stream.GetOutputStreamAt(0),
        InkPersistenceFormat.GifWithEmbeddedIsf);

    var bytes = new byte[stream.Size];
    var reader = new DataReader(stream.GetInputStreamAt(0));
    await reader.LoadAsync((uint)stream.Size);
    reader.ReadBytes(bytes);

    System.IO.Directory.CreateDirectory(System.IO.Path.GetDirectoryName(path));
    System.IO.File.WriteAllBytes(path, bytes);

    // Reload into the OS container to prove the strokes survived as ink. The WinUI
    // InkStrokeContainer has no public constructor, so use the OS one.
    var reloadStream = new InMemoryRandomAccessStream();
    var writer = new DataWriter(reloadStream.GetOutputStreamAt(0));
    writer.WriteBytes(System.IO.File.ReadAllBytes(path));
    await writer.StoreAsync();
    writer.DetachStream();

    var reloaded = new Windows.UI.Input.Inking.InkStrokeContainer();
    await reloaded.LoadAsync(reloadStream.GetInputStreamAt(0));

    return reloaded.GetStrokes().Count;
}

Attach a title bar and a page, and the two elements above are already a usable note app:

The Ink Notes sample app

The cream page and window chrome are ordinary WinUI styling around the canvas, left out of the snippets above.

The whole app

MainWindow.xaml
<?xml version="1.0" encoding="utf-8"?>
<Window
    x:Class="InkNotes.MainWindow"
    xmlns="http://schemas.microsoft.com/winfx/2006/xaml/presentation"
    xmlns:x="http://schemas.microsoft.com/winfx/2006/xaml"
    xmlns:muxc="using:Microsoft.UI.Xaml.Controls"
    Title="Ink Notes">

    <Grid x:Name="Root" Background="{ThemeResource ApplicationPageBackgroundThemeBrush}">
        <Grid.RowDefinitions>
            <RowDefinition Height="Auto" />
            <RowDefinition Height="*" />
        </Grid.RowDefinitions>

        <Border Grid.Row="0"
                Margin="24,12,24,0"
                Background="{ThemeResource CardBackgroundFillColorDefaultBrush}"
                BorderBrush="{ThemeResource CardStrokeColorDefaultBrush}"
                BorderThickness="1"
                CornerRadius="8"
                Padding="18,14">
            <StackPanel Spacing="12">
                <Grid>
                    <TextBlock Text="InkToolbar" FontSize="16" VerticalAlignment="Center" />
                    <Button x:Name="SaveButton"
                            Content="Save"
                            HorizontalAlignment="Right"
                            Click="OnSaveClick" />
                </Grid>
                <muxc:InkToolbar x:Name="Toolbar"
                                 TargetInkCanvas="{x:Bind Ink}"
                                 HorizontalAlignment="Left">
                    <muxc:InkToolbarBallpointPenButton x:Name="PenButton" />
                    <muxc:InkToolbarPencilButton />
                    <muxc:InkToolbarHighlighterButton />
                    <muxc:InkToolbarEraserButton />
                    <muxc:InkToolbarStencilButton x:Name="StencilButton" />
                </muxc:InkToolbar>
            </StackPanel>
        </Border>

        <Border Grid.Row="1"
                Margin="24,12,24,24"
                Background="#FFFDFCF8"
                BorderBrush="{ThemeResource CardStrokeColorDefaultBrush}"
                BorderThickness="1"
                CornerRadius="8">
            <muxc:InkCanvas x:Name="Ink" />
        </Border>
    </Grid>
</Window>
MainWindow.xaml.cs
// Copyright (c) Microsoft Corporation.
// Licensed under the MIT License.

using System;
using System.Threading.Tasks;
using Microsoft.UI.Xaml;
using Microsoft.UI.Xaml.Input;
using Windows.Storage.Streams;
using Windows.System;
using Windows.UI.Input.Inking;

namespace InkNotes
{
    /// <summary>
    /// A note page you can write on. The canvas and the toolbar do the work; this file only wires
    /// up the input devices and saving.
    /// </summary>
    public sealed partial class MainWindow : Window
    {
        public MainWindow()
        {
            this.InitializeComponent();

            // The canvas listens to pen only out of the box. Open it up to everyone.
            Ink.InkPresenter.InputDeviceTypes =
                Windows.UI.Core.CoreInputDeviceTypes.Touch |
                Windows.UI.Core.CoreInputDeviceTypes.Pen |
                Windows.UI.Core.CoreInputDeviceTypes.Mouse;

            // Ctrl+S saves too, for people who expect it.
            var save = new KeyboardAccelerator { Key = VirtualKey.S, Modifiers = VirtualKeyModifiers.Control };
            save.Invoked += async (_, args) =>
            {
                args.Handled = true;
                await SaveAsync();
            };
            Root.KeyboardAccelerators.Add(save);
        }

        private async void OnSaveClick(object sender, RoutedEventArgs e) => await SaveAsync();

        private static string NotePath => System.IO.Path.Combine(
            Environment.GetFolderPath(Environment.SpecialFolder.LocalApplicationData),
            "InkNotes",
            "note.gif");

        private async Task SaveAsync()
        {
            int strokes = await SaveAndReloadAsync(NotePath);
            Title = $"Ink Notes - saved {strokes} strokes";
            await Task.Delay(2000);
            Title = "Ink Notes";
        }

        /// <summary>
        /// GifWithEmbeddedIsf writes a normal GIF that any viewer can open, with the stroke data
        /// travelling inside it. Loading the file back returns editable strokes, not a flat picture.
        /// </summary>
        private async Task<int> SaveAndReloadAsync(string path)
        {
            var stream = new InMemoryRandomAccessStream();
            await Ink.InkPresenter.StrokeContainer.SaveAsync(
                stream.GetOutputStreamAt(0),
                InkPersistenceFormat.GifWithEmbeddedIsf);

            var bytes = new byte[stream.Size];
            var reader = new DataReader(stream.GetInputStreamAt(0));
            await reader.LoadAsync((uint)stream.Size);
            reader.ReadBytes(bytes);

            System.IO.Directory.CreateDirectory(System.IO.Path.GetDirectoryName(path));
            System.IO.File.WriteAllBytes(path, bytes);

            // Reload to prove the strokes survived as ink. The WinUI InkStrokeContainer has no
            // public constructor, so use the one from Windows.UI.Input.Inking.
            var reloadStream = new InMemoryRandomAccessStream();
            var writer = new DataWriter(reloadStream.GetOutputStreamAt(0));
            writer.WriteBytes(System.IO.File.ReadAllBytes(path));
            await writer.StoreAsync();
            writer.DetachStream();

            var reloaded = new InkStrokeContainer();
            await reloaded.LoadAsync(reloadStream.GetInputStreamAt(0));

            return reloaded.GetStrokes().Count;
        }
    }
}

Try it

Inking ships in the Windows App SDK experimental channel. Add Microsoft.WindowsAppSDK 2.5.4-experimental to a WinUI 3 project, add the two elements above, and you have somewhere to write.

💬 Tell us what you build

We want to hear about it: what you are building, what feels good, and anything that gets in your way. Bugs, rough edges and missing pieces are all useful to us, and the earlier they reach us the more we can do about them.

Start a discussion or file an issue at microsoft/microsoft-ui-xaml.

Happy inking 🙌

Team WinUI

Read the whole story
alvinashcraft
16 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Zero to Agent in 30 Minutes: Build Your First Agent with MCP

1 Share

Developers already have useful capabilities exposed through REST APIs. The Model Context Protocol (MCP) lets developers make those capabilities available to AI clients without rebuilding the underlying application.

In this episode of Zero to Agent in 30 Minutes, Bruce Hopkins, an AI developer, author, and longtime software educator, shows how to do that with an MCP server. His demo wraps an existing stock-data API in an MCP server so an MCP client can call it.

From REST API to MCP, step-by-step

  1. Start with an existing API. Identify the operations and data you want an AI client to access. The demo uses the Twelve Data API to retrieve current and historical stock prices.
  2. Create an MCP server. Use an MCP SDK to create the layer between the AI client and your existing application logic. In Bruce’s Python example, FastMCP handles the MCP interface while the stock-data functions remain separate.
  3. Expose capabilities as tools and resources. Register the operations the client should be able to discover and call. The stock-price data is exposed through MCP resources and tools that reuse the same underlying functions.
  4. Describe how the client should use them. Define clear names, inputs, descriptions, and prompts so the client understands what each capability does and what information it requires. Bruce’s example includes prompts for current prices, historical prices, and expected symbol and date formats.
  5. Connect the server to an MCP client. Run the server over a supported transport so the client can discover and call its tools and resources. 

You don’t need to replace the systems that already handle your application logic to help them work with agents. You can add an MCP interface around existing capabilities to give an AI client a standard way to discover and use them. Be sure to check out Bruce’s GitHub repo for working code you can adapt for your own APIs.

Coming next week

Next week, AI engineer Sajal Sharma returns to Zero to Agent in 30 Minutes to build a personal assistant on OpenClaw. He’ll show how an agent can keep tasks and notes in Markdown, run proactive automations, and deliver scheduled updates such as a regular morning briefing without waiting for a new prompt.

Follow along with Zero to Agent in 30 Minutes on Radar, or watch the latest episode on YouTube, Spotify, Apple, or wherever you get your podcasts. If you’re an O’Reilly member, you can watch live. Save your seat.



Read the whole story
alvinashcraft
3 hours ago
reply
Pennsylvania, USA
Share this story
Delete

ReviewBench: An open benchmark for AI code review

1 Share

Agentic code review is becoming an essential piece of how development happens. It helps you inspect pull requests, catch issues, and decide what deserves attention before code ships.

But the quality of existing AI reviewers can be hard to measure, and you need to know the strengths of a reviewer before you know if it will help you. Some reviewers surface more issues, some produce less noise, and some are stronger at catching critical problems while others surface smaller improvements, too. You may need code review to do different things within your workflow.

That makes it important to understand how reviewers actually compare: what different systems catch, what they miss, and the tradeoffs they make. A good code review benchmark should reflect the diversity of real pull requests, capture a broad set of review findings, and support meaningful breakdowns by severity, category, and precision-recall preferences. For teams building code review agents, the benchmark should also provide an offline signal that reliably tracks whether changes are likely to improve the experience in production. Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review, leaving a gap for a rigorous and reproducible evaluation methodology that brings these pieces together.

We built ReviewBench, a new code review offline benchmark, to address that gap, and it is available for you to use today. It follows the language, repo size, and size distribution of pull requests, modeled after over 100 million real pull requests on GitHub. It uses a multi-source golden set and a consistent evaluation rubric and has been independently validated by senior engineers. Just as important, with the help of ReviewBench, our offline evaluation of Copilot code review (CCR) has become more effective at anticipating the direction of production experiments, giving us greater confidence that measured improvements reflect meaningful gains for users.

In this post, we’ll walk through how ReviewBench is constructed, how it establishes reliable ground truth and scoring, and how to onboard your own code review system and submit results.

ReviewBench at a glance

1

What we built

A realistic, comprehensive benchmark for AI code review agents

103.9M

GitHub pull requests

Analyze distributions by language, repository size, and change shape.

Representative benchmark corpus

219 public pull requests across 19 languages, aligned to GitHub-wide distributions while preserving substantive review cases.

Multi-source golden set

  • Human reviewers
  • Frontier LLMs
  • Static analysis

Structured findings

Every finding is labeled for severity and category, enabling user-tailored slices.

Severity

  • Critical
  • Medium
  • Low

Category

  • Correctness
  • Security
  • Reliability
  • Maintainability
  • Testing
  • ......

Evaluation metrics

Four metrics measure both known and newly discovered issues.

  • Grounded precision
  • Grounded recall
  • Augmented precision
  • Augmented recall

Objective evaluation

Measure improvement and compare across agents objectively. Help users choose the reviewer that fits their needs the best.

2

How we keep it trustworthy

An auditable chain from rubric to expert validation and production checks

Published rubric

One explicit standard for all findings.

Human-labeled dev set

Senior engineers establish ground truth.

Calibrated grader

Aligned with human judgment.

Uniform labeling

Same standard across all sources.

Published agreement

Expert audit of benchmark quality.

Auditable end to end

96.6% agreement

Senior engineers independently labeled golden true-positives before release.

Offline signals that anticipate production

Benchmark movement is checked against online experiments.

  • Improvements tend to show up online
  • Regressions tend to show up online too

How ReviewBench works

Our benchmark is built around five principles:

1. Representative pull requests, not a demo set

We analyzed 103.9 million GitHub pull requests to characterize the real-world distribution of code review workloads. ReviewBench contains 219 pull requests from 187 public open source licensed repositories spanning 19 languages, with its language and repository-size distributions closely matching GitHub overall. The complete benchmark dataset is publicly available.

We make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most.

Quick corpus snapshot:

2. Broad ground truth discovery, independently judged

No single reviewer, whether human or model, can identify everything worth finding in a pull request. To build a broader and more reliable golden set for ground truth findings, we follow a three-stage process:

  • Gather candidate findings from diverse sources. We collect findings from real human reviewers, issues inferred from author follow-up commits, deterministic analysis tools, and multiple frontier LLMs across model families.
  • Semantically deduplicate overlapping findings. We merge findings that identify the same underlying issue, broadening coverage without allowing agreement across producers to artificially inflate the golden set or making it dependent on any one source’s blind spots.
  • Validate findings under a shared rubric. The source of a finding does not determine whether it is correct: a finding counts as a true positive only if it is true, relevant, and non-trivial. We use Claude Sonnet 5 as the LLM grader, applying a consistent evaluation rubric across all submissions. For transparency and reproducibility, we publish both the evaluation rubric and the judge used to apply it.

3. Metrics that measure both known and newly discovered issues

Most benchmarks report precision and recall against a fixed golden set. ReviewBench reports six metrics in two families:

  • Grounded precision, recall, and F1 score use only the existing gold-set labels. They provide the strict, apples-to-apples comparison: of the issues we already know about, how many did the agent find, and what share of its findings matched a known issue?
  • Augmented precision, recall and F1 score also evaluate findings that do not match anything in the golden set. The judge independently determines whether those unmatched findings are true or false positives, allowing a reviewer to receive credit for valid issues that no producer in the golden set surfaced

That distinction becomes more important as review agents become more capable. A fixed golden set inevitably becomes incomplete as systems discover issues its creators did not anticipate. Augmented metrics let ReviewBench recognize that behavior rather than automatically penalizing it. Because augmented recall expands the denominator based on what each agent discovers, we use grounded recall as the headline cross-system comparison and augmented metrics as an additional per-system diagnostic.

4. Configurable evaluation for different review preferences

There is no single universally optimal review experience. Some developers may want to focus only on critical issues, while others also value lower-severity, non-breaking findings. Some prefer broader coverage, while others prioritize precision and minimal noise. Others may have specialized needs, such as security- or privacy-focused review.

ReviewBench lets results be sliced by severity and category, while precision and recall capture different operating preferences. Users can also adjust β in the Fβ score to place more weight on recall for broader coverage or precision for lower noise. As these preferences change, the leaderboard is re-ranked accordingly, helping users identify the systems that best match their review priorities.

5. Internally audited and reproducibly evaluated

Before release, we asked senior engineers who had not participated in building the benchmark dataset to independently re-label every ground-truth finding from scratch. Their true/false-positive judgments agreed with ReviewBench 96.6% of the time. We version the benchmark dataset, judge, and matcher used in every evaluation, so results can be compared under the same benchmark configuration and revalidated when the benchmark changes. We also publish the validation methodology, agreement measurements, and known threats to validity, so readers can see how benchmark quality is assessed and where uncertainty remains.

Explore ReviewBench

ReviewBench’s research preview version is now available through the ReviewBench website, where you can explore the full benchmark, compare code review agents, and bring your own agent to evaluate and iterate.

With ReviewBench, you can:

  • Explore the full benchmark dataset. The complete ReviewBench dataset is publicly available, including the pull requests, findings, labels, severity and category annotations. This allows you to inspect exactly what systems are evaluated on and reproduce benchmark results.
  • Compare systems on the leaderboard. Results from evaluated code review agents using the full benchmark data are published on a common leaderboard, with views across overall performance, severity, category, and different precision–recall preferences.
  • Bring your own agent and hill-climb. The full benchmark dataset, evaluation methodology, LLM judge prompt, judge model configuration, and self-serve runner are publicly available, so you can evaluate your own code review agent, inspect its strengths and gaps, and iterate against the same benchmark configuration.

How we’ve used ReviewBench

We have used ReviewBench to evaluate Copilot code review (CCR) across successive iterations, giving us a consistent way to measure progress, catch regressions, and prioritize promising changes. Over time, this has helped us improve the product. One of the most valuable benefits of ReviewBench is that it provides an early offline signal of how a change to the product is likely to perform in production. Across experiments evaluated with ReviewBench before A/B testing, offline changes have consistently pointed in the same direction as what we see later in production.

A recent lite-tier experiment provides a concrete example of this broader pattern. We introduced a multi-model ensemble review that combines several independent model runs into a single review rather than relying on a single run. ReviewBench predicted higher precision, recall, and comment volume, along with lower cost per review.

To compare offline and production results, we use corresponding online signals. Addressed rate, our online counterpart to precision, is the percentage of CCR comments that an LLM determines prompted a developer to make a corresponding code change, based on the diff, thread, reactions, resolution state, and post-review code. For recall, we measure how much additional human review is still needed.

The online A/B test moved in the same direction as ReviewBench predicted: addressed rate (precision) rose 8.0%, recall rose 13.6%, and comment volume rose 61%, while cost per review fell 8.0%, all relative to the production control.

Comment volume alone, however, does not capture comment quality. More critical findings mean something very different from low-severity nits. ReviewBench’s severity-level evaluation captured this too: it predicted a 227% increase in critical comments, compared with 262% online, along with the same broader shift toward more moderate comments and fewer nits.

This gives us a fast and repeatable signal before running production experiments. Online experiments remain the ultimate measure of user impact, but ReviewBench gives us greater confidence in which changes are worth taking there.

How to submit your own run

  1. Sign in with GitHub on the ReviewBench website.
  2. Register your agent. Provide a container image, your configuration, and your own model key. We provide the judge.
  3. Try it on the test set. Run against a 25-PR test set with per-PR detail and repeat as you tune your configuration.
  4. Do a final run. When you’re ready, run the full set of 219 pull requests (three rounds), scored by the same judge as every other entry.
  5. Publish to the leaderboard. Your scores remain private until a maintainer reviews and approves the submission. Scores are published to the leaderboard only if they outperform the agent’s current leaderboard score, or if this is the agent’s first leaderboard entry.

We invite you to explore ReviewBench, evaluate your own system, challenge our assumptions, and help us improve the benchmark. We’re excited to collaborate with researchers and practitioners to make code review evaluation more open, reliable, and useful—and ultimately help move AI code review forward.

Acknowledgments

ReviewBench was a team effort across GitHub and Microsoft. We’re grateful to the researchers and engineers who built it: those who designed the methodology, curated the pull requests, built the golden set and the evaluation pipeline, and made the benchmark something anyone can run.

The post ReviewBench: An open benchmark for AI code review appeared first on The GitHub Blog.

Read the whole story
alvinashcraft
3 hours ago
reply
Pennsylvania, USA
Share this story
Delete

AI Agents Are Moving Into the Real World

1 Share
From: AIDailyBrief
Duration: 24:56
Views: 422

Meta is bringing Muse to smart glasses and a new wearable, while GrokBot is turning Teslas into voice-controlled personal assistants. As skeptics question whether consumers actually want AI agents, early users are finding real value in handing off life's annoying admin. In the headlines: Claude’s preliminary biology discovery, Trump’s “Super Intelligence” rebrand, and competing visions for AI governance at the UN.

The AI Daily Brief helps you understand the most important news and discussions in AI.
Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614
Get it ad free at http://patreon.com/aidailybrief
Learn more about the show https://aidailybrief.ai/

Read the whole story
alvinashcraft
3 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Ember 7.3 Released

1 Share

The Ember project is excited to announce the release of Ember v7.3. This is a standard minor release as part of the Ember Release Train process.

This release brings a new way to create reactive state that doesn't need a class, makes a serious dent in the size of the JavaScript bundle for apps that are on the modern build system, and fixes a handful of long-standing bugs in the router and elsewhere! 🎉😃

Ember.js 7.3

Ember.js 7.3 introduces one new feature per RFC #1071: tracked can now be used outside of classes, and, arguably more importantly, both forms of tracked let you configure equality so that setting a value to what it already was no longer triggers a re-render. The release also includes some internal restructuring that lets bundlers drop a lot more unused code from ember-source, and ships eight bugfixes. There are no new deprecations.

tracked() outside of classes

Since Ember Octane, the way you create reactive state in Ember has been to put a @tracked property on a class:

import Component from '@glimmer/component';
import { tracked } from '@glimmer/tracking';

export default class Counter extends Component {
  @tracked count = 0;

  increment = () => this.count++;
}

This works great and is still what we recommend for the vast majority of app code, but it does mean that if you want a single reactive value you first need a class to put it on. That gets in the way in a few places: helpers and modifiers that are written as plain functions, tests that want a bit of state to poke at, and demos where every extra line of boilerplate is a distraction from the thing you are actually trying to show.

Ember 7.3 implements RFC #1071, which does two things to the existing tracked import. First, when you call it as a function with an initial value it returns a standalone reactive value usable outside of classes:

import { tracked } from '@glimmer/tracking';

const count = tracked(0);
const increment = () => count.value++;

<template>
  Count is: {{count.value}}

  <button {{on "click" increment}}>add one</button>
</template>

Reading count.value in a template (or in a getter that a template uses) entangles with the render similarly to reading a @tracked property would, and writing to it causes a re-render. Alongside .value there are four function short-hands:

count.get();                // same as reading count.value
count.set(2);               // same as assigning count.value
count.update((n) => n + 1); // write based on the current value, without consuming it
count.freeze();             // prevent any further writes

One nice pattern that this unlocks is keeping mutable state private to a class while exposing a read-only view of it:

import { tracked } from '@glimmer/tracking';

export class Session {
  #user = tracked(null);

  get user() {
    return this.#user.value;
  }

  async login() {
    this.#user.value = await fetchCurrentUser();
  }
}

If any of this looks familiar it's because the idea has been floating around the ecosystem for a while. It was prototyped as Cell in Starbeam and has been available to Ember developers as cell from ember-resources. Now it's built in, with no extra import. Rather than @tracked being "magic, we can utilize tracked() as a storytelling tool to demystify how @tracked works.

Configurable equality

Until now, setting a @tracked property always notified consumers, even when you set it to the exact value it already held, so this.count = this.count re-rendered everything that read count. That was a historical choice and there was no way to change it. Both forms of tracked now accept an options object with an equals function that decides whether a write counts as a change.

The non-decorator form defaults to Object.is, so count.value = count.value does not re-render, and you can pass { equals: () => false } to get the old always-notify behaviour. The @tracked decorator keeps its always-notify default for backwards compatibility, and you can now opt a property into equality-based notification:

class Counter {
  @tracked({ equals: (a, b) => a === b }) count = 0;

  // this no longer causes a re-render
  noop = () => (this.count = this.count);
}

You can read more, including a couple of edge cases around passing plain objects as the initial value, in the API docs for tracked.

Introduced in emberjs/ember.js PR #21471

Smaller bundles for Vite apps

In the Ember 7.2 release blog we talked about setting type: "module" on ember-source and said that it wasn't enough on its own to shrink your bundle. This release is where that starts to pay off.

ember-source now declares "sideEffects": false in its package.json. That is a hint to bundlers like Vite and Rollup that importing one of Ember's internal modules never has side effects on any other module, which means the bundler is free to drop any module that your app never actually uses. Along with that, some of Ember's internals have been restructured so that a small app no longer pulls in the old rendering pipeline and some classic-era pieces.

The effect is significant. The compressed JavaScript for the hello-world app in Ember's own smoke tests went from around 64.5 kB before this work started to about 37 kB with these changes, or roughly 42% smaller. If you are curious about the details, the numbers are tracked in PR #21456.

This only affects apps that consume the ESM sources of ember-source directly, which is every app on the Embroider and Vite build system that has been the default since Ember 6.8. If you are still on the classic ember-cli build you won't see any change, which is one more reason to look at the Vite codemod if you haven't yet.

Introduced in emberjs/ember.js PR #21456 and PR #21462

Bug Fixes

Ember.js 7.3 introduces 8 bugfixes:

  • #21203 Fix @model becomes undefined or changes to the wrong route's model during Glimmer component willDestroy
  • #21591 Destroy dynamic modifiers that were set after the initial render, fixing a memory leak
  • #21409 Fix query params trigger model refresh unnecessarily
  • #21410 Fix query param redirects during active transitions
  • #21521 Treat nullish <LinkTo> @query as an empty query object
  • #21524 Allow CoreObject#init to be called with no arguments
  • #21406 Add a helpful assertion when {{component}} is given an unsupported argument
  • #21407 Improve {{debugger}} message for template-only components

A couple of these deserve a special mention. The first one fixes a bug that was reported in 2020 and, from the git history, has probably existed since @model was introduced in Ember 3.14. If a component in a route template read @model in its willDestroy hook while you were transitioning to a different route, it could see undefined or, worse, the other route's model. The fix adds a check on the controller's identity so the outlet cannot be redirected mid-teardown, and the new smoke test covers transitions to sibling, parent, cousin, and unrelated routes.

The second one is a memory leak that has been with us since Ember 3.25. If a dynamic modifier like {{this.mod}} started out as undefined and was set to a real modifier after the first render (or was swapped for a different modifier later), its destructor never ran when the element went away, so anything the modifier had set up, such as the floating-ui observers in ember-primitives, leaked. The fix registers the updating opcode with its block so teardown reaches it. The PR is tagged for backport to an LTS.

The query param fixes are part of a larger effort to improve the router's test coverage, and each of them closes an issue that people have been hitting for years: parent routes no longer re-run their model hook when you transition to a child route with unchanged query params, and redirecting from beforeModel back to the same route with different query params no longer loses those params or crashes on a direct visit. And if you have ever had a <LinkTo> blow up because you passed @query={{this.maybeParams}} and it happened to be null, that now works as expected.

Finally, a small security hardening that is not in the list above: the guard that stops set() from @ember/object walking through __proto__ and constructor in a path now also blocks prototype, closing a prototype pollution edge case. See PR #21451.

Documentation

The API docs for {{each}} and {{each-in}} now document that they support native Set and Map respectively, which they have done since the iterable refactor but never said so. See PR #21523. The Ember.Templates.helpers docs also link to @ember/helper correctly again (PR #21573).

Ember CLI 7.3

Ember CLI 7.3 is a maintenance release for ember-cli itself: no new features, no deprecations, and no bugfixes. The only changes there are the routine dependency updates that happen as part of the release train: ember-cli itself picked up small updates to things like testem and morgan, and the default app blueprint now generates apps on ember-source 7.3 with the current @embroider/macros, @glint/template, and eslint-plugin-warp-drive versions. None of those updates crossed a major version boundary, so there is nothing to do when you upgrade. As with 7.2, newly generated apps still get WarpDrive 5.8.

Tests run through testem directly

The @ember/app-blueprint that ember new uses did get one change worth knowing about. Since Ember 6.8, pnpm test in a new app has built the app with Vite and then handed the built output to ember test --path dist. In 7.3 that second step calls testem directly:

"test": "vite build --mode development && testem ci --port 0"

with a cwd: 'dist' line added to testem.cjs so testem serves the built app. Nothing changes about how your tests are written or which browser runs them. ember test was a thin wrapper around testem for Vite apps, and now it calls testem directly, removing a layer that could make it unclear which tool owned the flags you were passing. If you already have an app you do not need to change anything, but if you want the same setup you can copy the script and the one config line.

Introduced in ember-cli/ember-app-blueprint PR #307

Thank You!

A final note: this post was drafted with the help of artificial intelligence, and a human reviewed it in its entirety.

As a community-driven open-source project with an ambitious scope, each of these releases serves as a reminder that the Ember project would not have been possible without your continued support. We are extremely grateful to our contributors for their efforts.

Read the whole story
alvinashcraft
3 hours ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories