Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162319 stories
·
33 followers

Zero to Agent in 30 Minutes: Build Your First Agent with MCP

1 Share

Developers already have useful capabilities exposed through REST APIs. The Model Context Protocol (MCP) lets developers make those capabilities available to AI clients without rebuilding the underlying application.

In this episode of Zero to Agent in 30 Minutes, Bruce Hopkins, an AI developer, author, and longtime software educator, shows how to do that with an MCP server. His demo wraps an existing stock-data API in an MCP server so an MCP client can call it.

From REST API to MCP, step-by-step

  1. Start with an existing API. Identify the operations and data you want an AI client to access. The demo uses the Twelve Data API to retrieve current and historical stock prices.
  2. Create an MCP server. Use an MCP SDK to create the layer between the AI client and your existing application logic. In Bruce’s Python example, FastMCP handles the MCP interface while the stock-data functions remain separate.
  3. Expose capabilities as tools and resources. Register the operations the client should be able to discover and call. The stock-price data is exposed through MCP resources and tools that reuse the same underlying functions.
  4. Describe how the client should use them. Define clear names, inputs, descriptions, and prompts so the client understands what each capability does and what information it requires. Bruce’s example includes prompts for current prices, historical prices, and expected symbol and date formats.
  5. Connect the server to an MCP client. Run the server over a supported transport so the client can discover and call its tools and resources. 

You don’t need to replace the systems that already handle your application logic to help them work with agents. You can add an MCP interface around existing capabilities to give an AI client a standard way to discover and use them. Be sure to check out Bruce’s GitHub repo for working code you can adapt for your own APIs.

Coming next week

Next week, AI engineer Sajal Sharma returns to Zero to Agent in 30 Minutes to build a personal assistant on OpenClaw. He’ll show how an agent can keep tasks and notes in Markdown, run proactive automations, and deliver scheduled updates such as a regular morning briefing without waiting for a new prompt.

Follow along with Zero to Agent in 30 Minutes on Radar, or watch the latest episode on YouTube, Spotify, Apple, or wherever you get your podcasts. If you’re an O’Reilly member, you can watch live. Save your seat.



Read the whole story
alvinashcraft
50 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

ReviewBench: An open benchmark for AI code review

1 Share

Agentic code review is becoming an essential piece of how development happens. It helps you inspect pull requests, catch issues, and decide what deserves attention before code ships.

But the quality of existing AI reviewers can be hard to measure, and you need to know the strengths of a reviewer before you know if it will help you. Some reviewers surface more issues, some produce less noise, and some are stronger at catching critical problems while others surface smaller improvements, too. You may need code review to do different things within your workflow.

That makes it important to understand how reviewers actually compare: what different systems catch, what they miss, and the tradeoffs they make. A good code review benchmark should reflect the diversity of real pull requests, capture a broad set of review findings, and support meaningful breakdowns by severity, category, and precision-recall preferences. For teams building code review agents, the benchmark should also provide an offline signal that reliably tracks whether changes are likely to improve the experience in production. Existing benchmarks often make tradeoffs between label quality, coverage, and how well they represent real-world code review, leaving a gap for a rigorous and reproducible evaluation methodology that brings these pieces together.

We built ReviewBench, a new code review offline benchmark, to address that gap, and it is available for you to use today. It follows the language, repo size, and size distribution of pull requests, modeled after over 100 million real pull requests on GitHub. It uses a multi-source golden set and a consistent evaluation rubric and has been independently validated by senior engineers. Just as important, with the help of ReviewBench, our offline evaluation of Copilot code review (CCR) has become more effective at anticipating the direction of production experiments, giving us greater confidence that measured improvements reflect meaningful gains for users.

In this post, we’ll walk through how ReviewBench is constructed, how it establishes reliable ground truth and scoring, and how to onboard your own code review system and submit results.

ReviewBench at a glance

1

What we built

A realistic, comprehensive benchmark for AI code review agents

103.9M

GitHub pull requests

Analyze distributions by language, repository size, and change shape.

Representative benchmark corpus

219 public pull requests across 19 languages, aligned to GitHub-wide distributions while preserving substantive review cases.

Multi-source golden set

  • Human reviewers
  • Frontier LLMs
  • Static analysis

Structured findings

Every finding is labeled for severity and category, enabling user-tailored slices.

Severity

  • Critical
  • Medium
  • Low

Category

  • Correctness
  • Security
  • Reliability
  • Maintainability
  • Testing
  • ......

Evaluation metrics

Four metrics measure both known and newly discovered issues.

  • Grounded precision
  • Grounded recall
  • Augmented precision
  • Augmented recall

Objective evaluation

Measure improvement and compare across agents objectively. Help users choose the reviewer that fits their needs the best.

2

How we keep it trustworthy

An auditable chain from rubric to expert validation and production checks

Published rubric

One explicit standard for all findings.

Human-labeled dev set

Senior engineers establish ground truth.

Calibrated grader

Aligned with human judgment.

Uniform labeling

Same standard across all sources.

Published agreement

Expert audit of benchmark quality.

Auditable end to end

96.6% agreement

Senior engineers independently labeled golden true-positives before release.

Offline signals that anticipate production

Benchmark movement is checked against online experiments.

  • Improvements tend to show up online
  • Regressions tend to show up online too

How ReviewBench works

Our benchmark is built around five principles:

1. Representative pull requests, not a demo set

We analyzed 103.9 million GitHub pull requests to characterize the real-world distribution of code review workloads. ReviewBench contains 219 pull requests from 187 public open source licensed repositories spanning 19 languages, with its language and repository-size distributions closely matching GitHub overall. The complete benchmark dataset is publicly available.

We make one deliberate adjustment to this distribution: while language and repository size mirror GitHub directly, pull request size is weighted toward the reviewable middle and tail. This reduces the overrepresentation of tiny, single-file changes while preserving more substantive, multi-file pull requests where review quality matters most.

Quick corpus snapshot:

2. Broad ground truth discovery, independently judged

No single reviewer, whether human or model, can identify everything worth finding in a pull request. To build a broader and more reliable golden set for ground truth findings, we follow a three-stage process:

  • Gather candidate findings from diverse sources. We collect findings from real human reviewers, issues inferred from author follow-up commits, deterministic analysis tools, and multiple frontier LLMs across model families.
  • Semantically deduplicate overlapping findings. We merge findings that identify the same underlying issue, broadening coverage without allowing agreement across producers to artificially inflate the golden set or making it dependent on any one source’s blind spots.
  • Validate findings under a shared rubric. The source of a finding does not determine whether it is correct: a finding counts as a true positive only if it is true, relevant, and non-trivial. We use Claude Sonnet 5 as the LLM grader, applying a consistent evaluation rubric across all submissions. For transparency and reproducibility, we publish both the evaluation rubric and the judge used to apply it.

3. Metrics that measure both known and newly discovered issues

Most benchmarks report precision and recall against a fixed golden set. ReviewBench reports six metrics in two families:

  • Grounded precision, recall, and F1 score use only the existing gold-set labels. They provide the strict, apples-to-apples comparison: of the issues we already know about, how many did the agent find, and what share of its findings matched a known issue?
  • Augmented precision, recall and F1 score also evaluate findings that do not match anything in the golden set. The judge independently determines whether those unmatched findings are true or false positives, allowing a reviewer to receive credit for valid issues that no producer in the golden set surfaced

That distinction becomes more important as review agents become more capable. A fixed golden set inevitably becomes incomplete as systems discover issues its creators did not anticipate. Augmented metrics let ReviewBench recognize that behavior rather than automatically penalizing it. Because augmented recall expands the denominator based on what each agent discovers, we use grounded recall as the headline cross-system comparison and augmented metrics as an additional per-system diagnostic.

4. Configurable evaluation for different review preferences

There is no single universally optimal review experience. Some developers may want to focus only on critical issues, while others also value lower-severity, non-breaking findings. Some prefer broader coverage, while others prioritize precision and minimal noise. Others may have specialized needs, such as security- or privacy-focused review.

ReviewBench lets results be sliced by severity and category, while precision and recall capture different operating preferences. Users can also adjust β in the Fβ score to place more weight on recall for broader coverage or precision for lower noise. As these preferences change, the leaderboard is re-ranked accordingly, helping users identify the systems that best match their review priorities.

5. Internally audited and reproducibly evaluated

Before release, we asked senior engineers who had not participated in building the benchmark dataset to independently re-label every ground-truth finding from scratch. Their true/false-positive judgments agreed with ReviewBench 96.6% of the time. We version the benchmark dataset, judge, and matcher used in every evaluation, so results can be compared under the same benchmark configuration and revalidated when the benchmark changes. We also publish the validation methodology, agreement measurements, and known threats to validity, so readers can see how benchmark quality is assessed and where uncertainty remains.

Explore ReviewBench

ReviewBench’s research preview version is now available through the ReviewBench website, where you can explore the full benchmark, compare code review agents, and bring your own agent to evaluate and iterate.

With ReviewBench, you can:

  • Explore the full benchmark dataset. The complete ReviewBench dataset is publicly available, including the pull requests, findings, labels, severity and category annotations. This allows you to inspect exactly what systems are evaluated on and reproduce benchmark results.
  • Compare systems on the leaderboard. Results from evaluated code review agents using the full benchmark data are published on a common leaderboard, with views across overall performance, severity, category, and different precision–recall preferences.
  • Bring your own agent and hill-climb. The full benchmark dataset, evaluation methodology, LLM judge prompt, judge model configuration, and self-serve runner are publicly available, so you can evaluate your own code review agent, inspect its strengths and gaps, and iterate against the same benchmark configuration.

How we’ve used ReviewBench

We have used ReviewBench to evaluate Copilot code review (CCR) across successive iterations, giving us a consistent way to measure progress, catch regressions, and prioritize promising changes. Over time, this has helped us improve the product. One of the most valuable benefits of ReviewBench is that it provides an early offline signal of how a change to the product is likely to perform in production. Across experiments evaluated with ReviewBench before A/B testing, offline changes have consistently pointed in the same direction as what we see later in production.

A recent lite-tier experiment provides a concrete example of this broader pattern. We introduced a multi-model ensemble review that combines several independent model runs into a single review rather than relying on a single run. ReviewBench predicted higher precision, recall, and comment volume, along with lower cost per review.

To compare offline and production results, we use corresponding online signals. Addressed rate, our online counterpart to precision, is the percentage of CCR comments that an LLM determines prompted a developer to make a corresponding code change, based on the diff, thread, reactions, resolution state, and post-review code. For recall, we measure how much additional human review is still needed.

The online A/B test moved in the same direction as ReviewBench predicted: addressed rate (precision) rose 8.0%, recall rose 13.6%, and comment volume rose 61%, while cost per review fell 8.0%, all relative to the production control.

Comment volume alone, however, does not capture comment quality. More critical findings mean something very different from low-severity nits. ReviewBench’s severity-level evaluation captured this too: it predicted a 227% increase in critical comments, compared with 262% online, along with the same broader shift toward more moderate comments and fewer nits.

This gives us a fast and repeatable signal before running production experiments. Online experiments remain the ultimate measure of user impact, but ReviewBench gives us greater confidence in which changes are worth taking there.

How to submit your own run

  1. Sign in with GitHub on the ReviewBench website.
  2. Register your agent. Provide a container image, your configuration, and your own model key. We provide the judge.
  3. Try it on the test set. Run against a 25-PR test set with per-PR detail and repeat as you tune your configuration.
  4. Do a final run. When you’re ready, run the full set of 219 pull requests (three rounds), scored by the same judge as every other entry.
  5. Publish to the leaderboard. Your scores remain private until a maintainer reviews and approves the submission. Scores are published to the leaderboard only if they outperform the agent’s current leaderboard score, or if this is the agent’s first leaderboard entry.

We invite you to explore ReviewBench, evaluate your own system, challenge our assumptions, and help us improve the benchmark. We’re excited to collaborate with researchers and practitioners to make code review evaluation more open, reliable, and useful—and ultimately help move AI code review forward.

Acknowledgments

ReviewBench was a team effort across GitHub and Microsoft. We’re grateful to the researchers and engineers who built it: those who designed the methodology, curated the pull requests, built the golden set and the evaluation pipeline, and made the benchmark something anyone can run.

The post ReviewBench: An open benchmark for AI code review appeared first on The GitHub Blog.

Read the whole story
alvinashcraft
50 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

AI Agents Are Moving Into the Real World

1 Share
From: AIDailyBrief
Duration: 24:56
Views: 422

Meta is bringing Muse to smart glasses and a new wearable, while GrokBot is turning Teslas into voice-controlled personal assistants. As skeptics question whether consumers actually want AI agents, early users are finding real value in handing off life's annoying admin. In the headlines: Claude’s preliminary biology discovery, Trump’s “Super Intelligence” rebrand, and competing visions for AI governance at the UN.

The AI Daily Brief helps you understand the most important news and discussions in AI.
Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614
Get it ad free at http://patreon.com/aidailybrief
Learn more about the show https://aidailybrief.ai/

Read the whole story
alvinashcraft
50 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Ember 7.3 Released

1 Share

The Ember project is excited to announce the release of Ember v7.3. This is a standard minor release as part of the Ember Release Train process.

This release brings a new way to create reactive state that doesn't need a class, makes a serious dent in the size of the JavaScript bundle for apps that are on the modern build system, and fixes a handful of long-standing bugs in the router and elsewhere! 🎉😃

Ember.js 7.3

Ember.js 7.3 introduces one new feature per RFC #1071: tracked can now be used outside of classes, and, arguably more importantly, both forms of tracked let you configure equality so that setting a value to what it already was no longer triggers a re-render. The release also includes some internal restructuring that lets bundlers drop a lot more unused code from ember-source, and ships eight bugfixes. There are no new deprecations.

tracked() outside of classes

Since Ember Octane, the way you create reactive state in Ember has been to put a @tracked property on a class:

import Component from '@glimmer/component';
import { tracked } from '@glimmer/tracking';

export default class Counter extends Component {
  @tracked count = 0;

  increment = () => this.count++;
}

This works great and is still what we recommend for the vast majority of app code, but it does mean that if you want a single reactive value you first need a class to put it on. That gets in the way in a few places: helpers and modifiers that are written as plain functions, tests that want a bit of state to poke at, and demos where every extra line of boilerplate is a distraction from the thing you are actually trying to show.

Ember 7.3 implements RFC #1071, which does two things to the existing tracked import. First, when you call it as a function with an initial value it returns a standalone reactive value usable outside of classes:

import { tracked } from '@glimmer/tracking';

const count = tracked(0);
const increment = () => count.value++;

<template>
  Count is: {{count.value}}

  <button {{on "click" increment}}>add one</button>
</template>

Reading count.value in a template (or in a getter that a template uses) entangles with the render similarly to reading a @tracked property would, and writing to it causes a re-render. Alongside .value there are four function short-hands:

count.get();                // same as reading count.value
count.set(2);               // same as assigning count.value
count.update((n) => n + 1); // write based on the current value, without consuming it
count.freeze();             // prevent any further writes

One nice pattern that this unlocks is keeping mutable state private to a class while exposing a read-only view of it:

import { tracked } from '@glimmer/tracking';

export class Session {
  #user = tracked(null);

  get user() {
    return this.#user.value;
  }

  async login() {
    this.#user.value = await fetchCurrentUser();
  }
}

If any of this looks familiar it's because the idea has been floating around the ecosystem for a while. It was prototyped as Cell in Starbeam and has been available to Ember developers as cell from ember-resources. Now it's built in, with no extra import. Rather than @tracked being "magic, we can utilize tracked() as a storytelling tool to demystify how @tracked works.

Configurable equality

Until now, setting a @tracked property always notified consumers, even when you set it to the exact value it already held, so this.count = this.count re-rendered everything that read count. That was a historical choice and there was no way to change it. Both forms of tracked now accept an options object with an equals function that decides whether a write counts as a change.

The non-decorator form defaults to Object.is, so count.value = count.value does not re-render, and you can pass { equals: () => false } to get the old always-notify behaviour. The @tracked decorator keeps its always-notify default for backwards compatibility, and you can now opt a property into equality-based notification:

class Counter {
  @tracked({ equals: (a, b) => a === b }) count = 0;

  // this no longer causes a re-render
  noop = () => (this.count = this.count);
}

You can read more, including a couple of edge cases around passing plain objects as the initial value, in the API docs for tracked.

Introduced in emberjs/ember.js PR #21471

Smaller bundles for Vite apps

In the Ember 7.2 release blog we talked about setting type: "module" on ember-source and said that it wasn't enough on its own to shrink your bundle. This release is where that starts to pay off.

ember-source now declares "sideEffects": false in its package.json. That is a hint to bundlers like Vite and Rollup that importing one of Ember's internal modules never has side effects on any other module, which means the bundler is free to drop any module that your app never actually uses. Along with that, some of Ember's internals have been restructured so that a small app no longer pulls in the old rendering pipeline and some classic-era pieces.

The effect is significant. The compressed JavaScript for the hello-world app in Ember's own smoke tests went from around 64.5 kB before this work started to about 37 kB with these changes, or roughly 42% smaller. If you are curious about the details, the numbers are tracked in PR #21456.

This only affects apps that consume the ESM sources of ember-source directly, which is every app on the Embroider and Vite build system that has been the default since Ember 6.8. If you are still on the classic ember-cli build you won't see any change, which is one more reason to look at the Vite codemod if you haven't yet.

Introduced in emberjs/ember.js PR #21456 and PR #21462

Bug Fixes

Ember.js 7.3 introduces 8 bugfixes:

  • #21203 Fix @model becomes undefined or changes to the wrong route's model during Glimmer component willDestroy
  • #21591 Destroy dynamic modifiers that were set after the initial render, fixing a memory leak
  • #21409 Fix query params trigger model refresh unnecessarily
  • #21410 Fix query param redirects during active transitions
  • #21521 Treat nullish <LinkTo> @query as an empty query object
  • #21524 Allow CoreObject#init to be called with no arguments
  • #21406 Add a helpful assertion when {{component}} is given an unsupported argument
  • #21407 Improve {{debugger}} message for template-only components

A couple of these deserve a special mention. The first one fixes a bug that was reported in 2020 and, from the git history, has probably existed since @model was introduced in Ember 3.14. If a component in a route template read @model in its willDestroy hook while you were transitioning to a different route, it could see undefined or, worse, the other route's model. The fix adds a check on the controller's identity so the outlet cannot be redirected mid-teardown, and the new smoke test covers transitions to sibling, parent, cousin, and unrelated routes.

The second one is a memory leak that has been with us since Ember 3.25. If a dynamic modifier like {{this.mod}} started out as undefined and was set to a real modifier after the first render (or was swapped for a different modifier later), its destructor never ran when the element went away, so anything the modifier had set up, such as the floating-ui observers in ember-primitives, leaked. The fix registers the updating opcode with its block so teardown reaches it. The PR is tagged for backport to an LTS.

The query param fixes are part of a larger effort to improve the router's test coverage, and each of them closes an issue that people have been hitting for years: parent routes no longer re-run their model hook when you transition to a child route with unchanged query params, and redirecting from beforeModel back to the same route with different query params no longer loses those params or crashes on a direct visit. And if you have ever had a <LinkTo> blow up because you passed @query={{this.maybeParams}} and it happened to be null, that now works as expected.

Finally, a small security hardening that is not in the list above: the guard that stops set() from @ember/object walking through __proto__ and constructor in a path now also blocks prototype, closing a prototype pollution edge case. See PR #21451.

Documentation

The API docs for {{each}} and {{each-in}} now document that they support native Set and Map respectively, which they have done since the iterable refactor but never said so. See PR #21523. The Ember.Templates.helpers docs also link to @ember/helper correctly again (PR #21573).

Ember CLI 7.3

Ember CLI 7.3 is a maintenance release for ember-cli itself: no new features, no deprecations, and no bugfixes. The only changes there are the routine dependency updates that happen as part of the release train: ember-cli itself picked up small updates to things like testem and morgan, and the default app blueprint now generates apps on ember-source 7.3 with the current @embroider/macros, @glint/template, and eslint-plugin-warp-drive versions. None of those updates crossed a major version boundary, so there is nothing to do when you upgrade. As with 7.2, newly generated apps still get WarpDrive 5.8.

Tests run through testem directly

The @ember/app-blueprint that ember new uses did get one change worth knowing about. Since Ember 6.8, pnpm test in a new app has built the app with Vite and then handed the built output to ember test --path dist. In 7.3 that second step calls testem directly:

"test": "vite build --mode development && testem ci --port 0"

with a cwd: 'dist' line added to testem.cjs so testem serves the built app. Nothing changes about how your tests are written or which browser runs them. ember test was a thin wrapper around testem for Vite apps, and now it calls testem directly, removing a layer that could make it unclear which tool owned the flags you were passing. If you already have an app you do not need to change anything, but if you want the same setup you can copy the script and the one config line.

Introduced in ember-cli/ember-app-blueprint PR #307

Thank You!

A final note: this post was drafted with the help of artificial intelligence, and a human reviewed it in its entirety.

As a community-driven open-source project with an ambitious scope, each of these releases serves as a reminder that the Ember project would not have been possible without your continued support. We are extremely grateful to our contributors for their efforts.

Read the whole story
alvinashcraft
50 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

AWS Weekly Roundup: Amazon Bedrock Managed Agents powered by OpenAI, Q3 service availability updates, Kiro workflows, and more (October 5, 2026)

1 Share

Last week, we announced a public preview of Amazon Bedrock Managed Agents powered by OpenAI, built on a customized version of OpenAI’s Agents API engineered to be AWS-native and integrated with AWS resources. You can now build agents optimized for OpenAI models that run entirely inside AWS with the identities, permissions, and governance controls you already use.

You can choose an execution environment: self-hosted compute to use an existing development machine, container, or compute environment or Amazon Bedrock AgentCore Runtime, for managed runtime sessions and configurable storage in your AWS account. To learn more, visit the Amazon Bedrock documentation.

In addition, we are adding new frontier models on Amazon Bedrock to expand your model choices:

  • OpenAI GPT-6.1 Sol: An upgrade to GPT-6 Sol, GPT-6.1 Sol delivers exceptionally strong performance on agentic coding, computer use, and professional work. According to OpenAI, it approaches GPT-6 Astra across demanding evaluations at roughly one-fifth of the cost, giving developers more room to build and run capable agents at scale. To learn more, visit the GPT-6.1 Sol model card.
  • OpenAI GPT-6 Astra UltraFast mode: Ultrafast is a premium speed tier for GPT-6 Astra, built for workloads where speed matters most. According to OpenAI, Ultrafast delivers up to 6x faster inference in the API, with up to 300 tokens per second. The Amazon Bedrock inference engine delivers the performance, security, and reliability required for production workloads. To learn more, visit the GPT-6 Astra model card.
  • Anthropic Claude Sonnet 5.5: Claude Sonnet 5.5 is a smarter, more efficient Sonnet and a step up from Sonnet 5, making it a natural upgrade for teams already building on Sonnet. It’s stronger for coding, completing well-scoped tasks as part of a larger coding strategy such as building and fixing features with Claude in the same session or verifying output against requirements. To learn more, visit the Claude Sonnet 5.5 model card.
  • SpaceXAI Grok 4.7: Grok 4.7 builds on Grok 4.6 with better mixed-document handling, more dependable repo-scale coding with planning and error recovery, and enhanced browser-use agents for form fills and portal navigation. To learn more, visit the Grok 4.7 model card.

Last week’s launches
Here are some launches that got my attention:

  • AWS Well-Architected Agent (preview): You can use an AI-powered agent service that analyzes your AWS environment to deliver targeted, contextual recommendations for improving your applications’ cost, security, performance, and resilience. The agent analyzes your infrastructure, understands your unique business goals, and delivers contextual recommendations.
  • Amazon Aurora PostgreSQL supports direct querying of Apache Iceberg and Parquet data: You can directly query operational data together with data stored in data lakes in Apache Iceberg and Parquet formats using your existing PostgreSQL applications and tools, without extract, transform, and load (ETL) pipelines or data duplication.
  • Amazon S3 Tables support all Apache Iceberg V3 data types: Amazon S3 Tables add support for geometry, geography, unknown, and nanosecond timestamp data types, along with column default values, as defined in the Iceberg V3 specification. You can now store geospatial coordinates and nanosecond-precision event times natively instead of encoding them in strings or integers.

For a full list of AWS announcements, be sure to keep an eye on the What’s New with AWS page.

AWS service availability updates
When the availability of an AWS service or feature changes, we provide customers guidance in AWS Product Lifecycle Changes on available alternatives and support for migration so that disruptions to your operations are minimized. The following lifecycle changes were updated on September 29, 2026.

Services moving to Maintenance (no longer accessible to new customers starting October 29, 2026):

Services entering Sunset:

Services reaching End of Support (as of September 29, 2026):

  • Amazon Mechanical Turk

We understand that changes in availability can impact your operations. For specific guidance, consult the relevant service documentation or contact AWS Support.

Other AWS news
Here are some additional projects and news items you may find interesting:

  • Introducing Kiro workflows: Kiro workflows enable you to carry out complex tasks from start to finish with multiple agents and less supervision. We’ve been building Kiro itself with workflows, including the new cloud configuration, cloud sessions, and most of the workflow experience.
  • Introducing Strands Decider: Strands Decider is one of a new class of decision models or system one models, a type of model that has been gaining significant attention since TypeSafe AI’s launch of Jev earlier this month. Strands Decider 2B is a small, open source, decision model optimized for fast experimentation, local development, and innovation.
  • New FDE pathways for AWS Partners: On June 30, AWS announced the Forward Deployed Engineering (FDE) organization, backed by a $1 billion investment, and extended this hands-on delivery approach to AWS Partners through the Partner-Led FDE motion. Now, AWS Partners have a structured way to build and validate that depth with three new Partner FDE pathways and credentials that recognize the applied proficiency required to deliver production agentic AI.

For a full list of AWS blog posts, be sure to keep an eye on the AWS Blogs page.

Learn more about AWS, browse and join upcoming AWS-led in-person and virtual events, startup events, and developer-focused events, including upcoming AWS re:Invent and AWS Community Days. Join the AWS Builder Center to connect with builders, share solutions, and access content that supports your development.

That is all for this week. Check back next Monday for another Weekly Roundup!

— Channy

Read the whole story
alvinashcraft
50 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Tray-tober Day 5: Tray Do

1 Share

The most classic programming project is always a to-do list, and Tray-tober is no exception with Tray Do. It is a todo list, task tracker, always a click away in your tray. One little delighter is there is a ‘hit list’ options where tasks can be singled out to be prioritized on a separate list. This means the tray icon will change to reflect how many more tasks remain on your list keeping you focused.

What does Tray Do not have?

  • No web view
  • No login
  • No gamification mechanisms
  • No dark patterns to keep you using the app

Let small jobs have small solutions not aiming to get you pulled deeper into their world.

Happy Tray-tober!

Joe

Read the whole story
alvinashcraft
50 minutes ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories