Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162029 stories
·
33 followers

Daily Reading List – September 30, 2026 (#878)

1 Share

Last night, I went to the San Diego Padres playoff game against the Chicago Cubs. It was the rowdiest, loudest, most fun baseball game I’ve attended in my life. The downside? I’m very hoarse and had two major speaking gigs today. So yeah, I plan well.

[blog] Gemini 4 Argon: our next era of frontier intelligence. Are you not entertained? We shared information about this model. Numbers look … great.

[article] The best AI leaders are ‘a little bit off the wall,’ says Google Cloud executive. I did say that. Here’s a summary of some of the points I made in Boston a couple weeks back.

[blog] Build an agentic software factory, starting with one bug. As you build up a set of AI agents to help you with software, do it incrementally. Useful lesson here.

[blog] Graph Workflows in ADK: Everything You Need to Know. I liked this explanation of working our way up to more sophisticated agent architectures.

[blog] State of agent skills. Vercel sits on a lot of data here thanks to running skills.sh. Read this for a look at what skills people are installing, for which industries, and how often.

[blog] State of Markets II. We’re getting a bunch of “state of XX” things lately. The a16z folks call out a bunch of data about companies using AI, and more.

[blog] Vulnerability Discovery and Exploitation Trends in the AI Era. Sobering data about the rise in vulnerabilities, and corresponding increase in exploitation. At least there are suggestions for what to do next!

[blog] Voice Agents Can Just Do Things. I buy this more than I did a year ago. But I also don’t want this to be the ONLY interface. Sometimes it’s better to type things out, or I’m in a space where I don’t want to be barking commands to a robot.

[blog] Data Agent Kit is now GA: Bring Google Data Cloud to any coding agent. Do some fairly sophisticated analytics and data science work in whatever AI coding tool you prefer.

Want to get this update sent to you every day? Subscribe to my RSS feed or subscribe via email below:



Read the whole story
alvinashcraft
52 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Your app’s frontend UI is now optional

1 Share

I have 334 apps on my phone. You probably have more than me. My browser bookmarks and history are stuffed with sites I infrequently use. We’re inundated with “apps” that require us to learn some bespoke interface just to get something done. That era is coming to an end. Quickly.

Don’t take my word for it. Many are acknowledging that AI agents have made many apps unnecessary. Oh, I still need the data or function from that app. But I never want to “see” it again. Drop the user interface. Go headless. Why am I still logging into a system three times a year to mark my vacation time? Or navigating multiple travel sites to find the best deals? Or even doing “regular” web searches to find stats to include a presentation? I just want my agent, harness, or super-app (like Gemini, Grok Bot, Meta Muse, Claude Cowork) to do it.

So if we’re not building a bunch of static frontends, what should developers do instead? Here are a few suggestions. Let’s look at two categories of tools that developers should create. And then three types of activities developers (or agents) can do with those tools.

Create CLIs, APIs, skills, and MCP servers for anywhere access to data and functionality

Good models have “computer use” where they can navigate your web frontend if they need to. But spelunking your DOM wastes tokens and time. At least create some WebMCP tools on your site to give agents a better shot at completing a task.

But honestly, those harnesses, super-apps, and agents just want your data and commands. Give it to them! Salesforce announced this headless experience in a partnership with Anthropic (ClaudeForce). Box too. HubSpot as well. Expect most SaaS products to do this. Same with platform companies. Google shipped an Android CLI for use in any harness, 150 agent skills, and tons of remote managed MCP servers. UI optional.

We did before, but now we have even more options to build these types of tools. There are toolkits for creating MCP servers, generating skills, and even cranking out a CLI. Help out your users by giving them the tools they need to get to your data and functionality without forcing them to use your UI.

Create A2UI or MCP Apps components for dynamic rendering

The technology is now there to create more personalized and interactive experiences for humans.

A2UI is a protocol that enables agents to create dynamic user interfaces. You have the data and you have the available widgets that an agent sends a client to render. You’re not shipping arbitrary HTML and JavaScript; the agent sends component descriptions that get safely rendered client side. There are renderers for Angular, React, Flutter, and more. Developers should consider building components that agents can use to compose a beautiful, personalized site for the user.

Then you have MCP Apps. Have you looked at these? Here you’re returning interactive HTML interfaces that get rendered in your agentic chat experience. Developers can build these so that chat users get more than boring text results in their sessions.

Build these portable UI components that users can apply to get richer experiences on whatever surface they can render them on.

Retrieve info or trigger action from where I’m at

Let’s now look at the user.

Maybe I don’t want to use your fancy portal to get my work done. Give me those CLIs, agents, MCPs, or MCP Apps to retrieve the data I want, in the surfaces I’m already using.

Here in my Google Antigravity CLI, I’ve got a reference to the remote MCP Server for Cloud Run. I could also use a locally installed CLI to do the same things.

Using an existing MCP or CLI, I can do most anything with Google Cloud Run. Want a list of running services? Don’t break flow and bounce to a fixed UI somewhere else. Just get the data in your harness or super-app of choice.

Using those tools created by app or product owners, we can now aggregate or interact with systems however we want. Maybe instead of a web front end, you build an agent. Cool. I took a hotel concierge agent that might have warranted a whole fancy website, and used it from within Gemini Enterprise instead.

Need something more interactive? I built an MCP App that talked to Google Cloud Run and returned a fancy dashboard that I can use from within Gemini Enterprise. Why leave my super-app if I don’t need to? First I configured my MCP App (running in Cloud Run) as a data source.

And now anyone can chat with it, getting a full UI back in Gemini Enterprise chat experience.

Build on-the-fly visualizers, apps, and pages personalized to me

The era of personal software is here. Maybe that software is disposable, or just for you. Possibly for a small circle of friends. Whatever. Build entirely new interfaces that work exactly for you by mashing up the APIs, CLIs, and MCPs available to you.

For example, I could use a series of pre-built A2UI components to create a hotel website that reacts to the situation at hand. Completely new user? Show one thing. Checked-in hotel guest, here’s another. Frequent guest coming to book a room? Here you go. Instead of building dozens of static pages to anticipate each scenario, I have ONE PAGE that applies relevant components on the fly.

Build wherever-I-want-it experiences to get work done

This is a variation of the previous two. You might want to stay in your tool of choice, but also want a customized interface just for you. Let’s try that.

We released a new MCP server for the Google Cloud CLI. It executes gcloud CLI commands remotely. This is handy if you’re on an agentic client that can’t execute CLIs, trusted or otherwise.

I could use this to build a full Google Cloud management experience as a Chrome plugin! I can’t easily execute CLI commands in the Chrome sandbox without some hoops. No problem. I can use the logged in user to invoke the MCP server via agent inside a custom plugin. This makes it possible to do almost anything in Google Cloud. from anywhere.

In this example, I spun up a Google Cloud Pub/Sub topic while browsing my daily reading list.

Look, static frontends won’t become truly optional for a while. But developers, don’t wait too long to adapt. These personal agents are growing fast, and fairly soon, many websites and mobile apps will have significantly more agent visitors than human ones. Build accordingly!



Read the whole story
alvinashcraft
57 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Updates to Copilot Code Reviews for Azure Repos

1 Share

We announced the public preview of Copilot Code Review for Azure Repos in late August. Since then, we have continued to make improvements, fix bugs, and add new features.

We wanted to share an update on some of the recent changes.

Effort levels

The effort level controls how much time Copilot spends reviewing a pull request. This gives you more control over the depth of the review and can help reduce costs when a detailed code review is not needed.

screen of effort level

You can set the default effort level to Lite or Balanced at the project and repository level. Project admins can set a default for all repositories in a project, while individual repositories can set their own default when needed.

screen shot of effort level settings

This gives teams the flexibility to choose the right level of review based on the needs of each project or repository.

Better way to resolve comments

We’ve made it easier to resolve comments as you apply suggestions.

Previously, you had to apply the suggested changes, commit them, and then go back and resolve each comment individually. Now, when you apply and commit suggested changes from a Copilot review, you can resolve all related comments at the same time.

image of resolving suggestions during commit

This makes it faster to work through pull request comments and gives you a clearer view of what still needs your attention.

Audit log events

Copilot Code Review usage is charged to the subscription for each code review, making it important for organization administrators to understand where Copilot Code Review is enabled and if any settings have been changed.

To provide better visibility, we have added several new events to Azure DevOps audit logs. These events capture:

  • Copilot Code Review is enabled or disabled at the organization, project, or repository level.
  • Changes to the configured agent pool
  • Changes to custom instructions at the organization or project level
  • Changes to the default effort level at the project or repository level
  • Branch policies added for automatic code reviews at the project or repository level
  • List item

These events are available on the Auditing page in Azure DevOps, where administrators can review Copilot Code Review configuration changes across the organization.

audit log image

More features to come

These features are currently rolling out to customers and should be available to everyone over the next two to three weeks.

We will continue to address customer feedback and add new features over the next couple of months.

If you have not already, give Copilot Code Reviews for Azure Repos a try and send us your feedback. We would love to hear what you think.

The post Updates to Copilot Code Reviews for Azure Repos appeared first on Azure DevOps Blog.

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete

Claude for Government is now generally available

1 Share
Claude for Government is now generally available
Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete

Evaluating AI-Generated Frontend Code: What Should We Actually Test?

1 Share

AI can now generate a surprising amount of frontend code from a short description. A developer can ask for a form, a table, a modal, a settings page, or a dashboard view and get something that looks usable almost immediately. It may compile, render, and even arrive with a few tests. That is useful, but it also creates a problem: The first version of the UI can look more complete than it really is.

Frontend code is often judged too quickly. If the build passes and the screen looks close to the design, it is tempting to treat the generated code as mostly done. But a user interface is not just a collection of components on a page. It is a path someone has to move through. It has to handle input, state, errors, loading, navigation, focus, responsiveness, and accessibility. Some of the most important failures are not visible in a screenshot.

This is why teams need better ways to evaluate AI-generated frontend code. They need evidence that the UI is ready for people to use.

Build and render are only the starting line

The easiest checks are usually the first ones teams run. Does the code compile? Does the page render? Are there obvious console errors? Does the component appear in the browser? Those checks matter, but they are only the starting line. A page can render while the form is difficult to complete. A modal can appear while focus remains behind it. A generated test can pass while the actual user flow is broken.

This is especially important with AI-generated code because the output often has a polished surface. The code may be formatted well, the component names may sound reasonable, and the test file may make the change look more complete than it is. That polish can make reviewers less likely to slow down and ask whether the interface actually works. A better evaluation process starts with a simple assumption: generated frontend code is a draft until the user behavior has been checked.

Start with the structure of the page

Before looking at more complex behavior, it is worth checking whether the generated UI has a sound structure. Frontend evaluation should include the basic semantics of the page, not just the visual layout.

That means checking whether the code uses native HTML where possible. A button should usually be a button, not a clickable div. A link should be used for navigation, not for actions that behave like buttons. A form field should have a label that is connected to it. These details are easy to overlook because the UI may look fine without them, but they affect how people navigate, how assistive technologies interpret the page, and how maintainable the code will be later.

AI tools sometimes choose generic containers where native elements would be better. They may also add ARIA without using it correctly. ARIA stands for Accessible Rich Internet Applications, a set of attributes defined by the W3C to help make web interfaces more accessible when native HTML is not enough. The W3C’s WAI-ARIA overview is a useful reference. ARIA can be important, but it should not be used as a substitute for the right HTML element. The first evaluation question should be: Did the generated code use the right building blocks?

Check the keyboard path

A useful frontend evaluation should include the keyboard path through the interface. Many users rely on keyboards or keyboard-like navigation, and keyboard testing also exposes problems in the interaction model.

The simplest test is often the most revealing: put the mouse aside and try to complete the task. If the flow becomes confusing, the generated code is not ready. Can you reach the important controls? Is the focus order logical? Can you open and close a dialog without a mouse? When the dialog closes, does focus return to a sensible place?

These checks are especially important for generated UI because AI can produce interactions that work for the most obvious mouse path but fail in less visible ways. A custom dropdown, for example, may open on click and look finished in a demo, but it may not respond correctly to keyboard input. That is not a small edge case. It is part of whether the interface is usable.

Test focus, not just clicks

Click-based tests are useful, but they can hide important problems. A test that clicks a button and waits for a success message may pass even when the same flow is frustrating for someone who is navigating by keyboard.

Focus behavior deserves its own attention, especially when the UI changes after the user takes an action. For example, when a form submission fails, the user should not be left guessing what happened. The error should be visible, connected to the relevant field when appropriate, and reachable in a way that makes recovery clear. In many cases, focus should move to the first error or to a summary that explains what needs attention.

The same idea applies to modals. When a modal opens, focus should move into it. When it closes, focus should return to the element that opened it. These are small details in code, but they make a large difference in whether the UI feels predictable.

Evaluate what happens when things go wrong

Generated frontend code often looks best in the happy path. The user fills everything in correctly, the network responds quickly, the data shape is exactly as expected, and nothing fails. Real interfaces spend a lot of time outside that path.

A practical evaluation should check what happens when data is missing, delayed, empty, invalid, or returned in an unexpected state. This is what I mean by loading, error, and empty states. They are the parts of the interface that explain what is happening when the ideal path breaks down. A loading state should help the user understand that something is in progress. An error state should explain what went wrong and what the user can do next. An empty state should make it clear whether there is nothing to show, whether the user needs to take action, or whether something failed quietly.

These cases are easy to leave for later because the happy path is usually enough to make the screen look finished. But users will eventually hit the less perfect paths. A generated component may include a spinner because the prompt asked for one, but that does not mean the loading experience is useful. An error message may say “Something went wrong,” but offer no recovery. Evaluation should include these cases because this is where many real user experiences break.

Test the full user flow

Component-level checks are helpful, but they do not always tell the full story. A component can work by itself and still fail when it is placed inside a larger flow.

That is why AI-generated frontend code should be evaluated through user tasks. Can someone start the flow, understand what is expected, recover from a mistake, submit successfully, and see what changed afterward? Does the interface still work on a smaller screen? Does the state remain consistent if the user goes back, edits something, or retries after a failure?

This is where Playwright-style tests or other end-to-end tests can be useful. The goal is not to automate every possible interaction. The goal is to protect the flows that matter most. A good test should determine whether the user can complete the task the component is supposed to support.

Use accessibility checks, but do not stop there

Automated accessibility checks are useful and should be part of the evaluation process. They can catch missing labels, invalid ARIA usage, some contrast issues, landmark problems, and other common mistakes. They are especially helpful when AI-generated code is moving quickly because they catch issues before they become repeated patterns.

But automated checks are not a complete accessibility review. They cannot fully judge whether a flow is understandable, whether focus movement feels natural, or whether instructions are clear. Passing an automated accessibility scan does not mean the UI is accessible. It means some common problems were not detected.

The best approach is to combine automated checks with behavior-based review. Run the tools, but also use the interface. Navigate by keyboard. Trigger an error. Try the empty state. Look at the generated code and ask whether native HTML could do more of the work. Accessibility evaluation is strongest when it is part of normal frontend quality, not a separate pass at the end.

Review the generated tests too

When AI generates code, it may also generate tests. That sounds helpful, but those tests need to be reviewed with the same care as the code.

Generated tests often reflect what the implementation already does. They may check that text appears, that a function was called, or that a component was rendered. Those checks are not useless, but they can create false confidence if they do not test meaningful behavior. A better review asks what the tests would catch if the UI broke. Would they fail if a validation error was unclear? Would they fail if the retry button did not work? Would they fail if keyboard navigation was broken?

If the answer is no, the tests may be documenting the implementation more than protecting the user experience. Teams can use AI to help write better tests, but the prompt matters. “Write tests for this component” is too vague. A better request explains the behavior that matters, such as validation recovery, loading behavior, successful submission, and focus movement. Even then, the generated tests still need human review.

Decide what evidence is enough

Not every UI change needs the same level of evaluation. A small copy update does not require the same review as a new checkout flow, onboarding flow, or account settings page. Teams need judgment.

A useful approach is to match the evaluation to the risk of the change. If the generated code affects a critical user flow, collects user input, changes navigation, introduces a custom interaction, or handles important status messages, it deserves deeper testing. If it reuses stable components in a familiar pattern, the review may be lighter.

There is no need to create a checklist for every pull request; clarity on what evidence is enough suffices. For some changes, a quick review and component test may be fine. For others, the team should expect keyboard testing, accessibility checks, error-state review, and a user-flow test. The point is to avoid treating all generated code as equally trustworthy just because it looks polished.

Human review still matters

AI can generate code and suggest tests, but it cannot fully understand the product, the users, or the trade-offs behind a frontend decision. It does not know which flows are most important, which interaction patterns users already rely on, or where inconsistency will cause confusion.

That is why human review remains central. The reviewer’s role is to ask whether the generated solution fits the system and supports the user’s task. Sometimes that means accepting the generated code. Sometimes it means asking for a simpler native element, reusing an existing component, improving the error recovery, or adding a test that reflects real behavior.

The more code AI generates, the more important this judgment becomes.

What should we actually test?

When AI writes frontend code, teams should test the parts of the interface that users depend on. That includes structure, keyboard access, focus behavior, loading and error states, form validation, responsive behavior, accessibility checks, and full user-flow completion. It also includes reviewing the generated tests themselves to make sure they protect behavior rather than merely confirming the current implementation.

The goal is not to slow down AI-assisted development. The goal is to make it safer to use. If AI reduces the time spent producing a first draft, teams have an opportunity to spend more time asking whether the software actually works properly.

That may be the real shift. In frontend development, the value of AI is not just faster code. It is the chance to move more engineering attention toward evaluation, user behavior, and quality.

AI-generated UI should not be trusted because it looks complete. It should be trusted because the team has checked the right things.

AI use acknowledgment

AI assistance was used lightly for phrasing, editing, and tightening parts of this draft. The article’s ideas, structure, examples, and final review are my own.

Author’s note

The views expressed are my own and do not represent those of my employer.



Read the whole story
alvinashcraft
15 hours ago
reply
Pennsylvania, USA
Share this story
Delete

The Night Sky Is Getting 10% Brighter Every Year

2 Shares
fjo3 quotes a report from The Guardian, written by Tove Danovich: Every year, the night sky is becoming 10% brighter. Eighty percent of people live under light-polluted skies. And while the UN declared a healthy environment -- clean air, clean water -- to be a human right, I believe we also all have a right to darkness as well. This doesn't mean turning off every streetlamp. But do we really need bright white floodlights at the front of every garage? Replacing lights with ones that are bright enough to do the job -- and no brighter -- might make it so I wouldn't have to get in my car and drive for hours to see the night sky as it actually shines above us. To preserve the darkness would mean rules and enforcement around the warmth of the lights, where they point, and how bright they can be. It would mean not putting in three lights when one would do. If we recognized that darkness is something we need, we'd become more careful about chasing it away. [...] Access to darkness is about more than stargazing. The night is its own habitat. Artificial lights confuse animals who use the moon to navigate, whether they're baby sea turtles navigating toward a parking lot instead of the ocean or moths who fly in circles around a lightbulb. Lightning bugs and frogs need darkness to complete their courtship rituals. Migrating birds get thrown off course by the bright lights of big cities. But the effects of light pollution linger even after it's daylight again, becoming visible. Research has shown light pollution is breaking the relationship between plants and pollinators, changing the time of year trees break into flower, and disrupting circadian rhythms for all living creatures. This includes humans. Whether it's artificial lights indoors or creeping in through the window at night, studies have found that we need darkness in order to rest, recover and stay healthy.

Read more of this story at Slashdot.

Read the whole story
alvinashcraft
15 hours ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories