Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
161836 stories
·
33 followers

Week in Review: Most popular stories on GeekWire for the week of Sept. 20, 2026

1 Share

Get caught up on the latest technology and startup news from the past week. Here are the most popular stories on GeekWire for the week of Sept. 20, 2026.

Sign up to receive these updates every Sunday in your inbox by subscribing to our GeekWire Weekly email newsletter.

Most popular stories on GeekWire

Read the whole story
alvinashcraft
32 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Android Weekly Issue #746

1 Share
Articles & Tutorials
Sponsored
Debugging mobile apps is weird: intermittent connections, mid-onboarding drop-offs, edge cases on devices you've never tested. bitdrift captures 100% of data, unsampled and in real time, so it’s immediately queryable by engineers and agents. Try bitdrift: mobile observability for the real world.
alt
Enrique López-Mañas examines layered defenses against SDK tampering: obfuscation, request signing, device attestation, and replay protection.
Arnaud Giuliani explains how Koin Compiler 1.2 extends compile-time dependency validation to all Koin DSL styles.
Sponsored
Android apps face over 14 million attacks annually, and most vulnerabilities are found too late. DerScanner brings enterprise-grade SAST to your APK: false positives cleared, context-aware fixes ready to apply. 14-day community edition access.
alt
Marcin Moskala explains why a data class beats a sealed class for representing loading, refresh, and error states.
Nirmal Jeffrey explains how sharing Kotlin logic across Android, iOS, and web kept his app's rules in sync.
Kevin Desai builds a custom Compose layout recreating Libby's herringbone book arrangement using geometry and placeWithLayer.
Akshay Nandwana explains how PocketCommunity turns Gemini-generated component trees into native Compose UI via A2UI.
Jan Rabe explains routing emulator HTTPS traffic through a local proxy to reach a Dockerized dev stack.
Nav Singh recounts how a forgotten feature flag left production on a legacy path, exposing monitoring gaps.
Andrew Malitchuk builds a Kotlin Multiplatform OCR SDK unifying Tesseract and Vision into one structured document model.
John O'Reilly demonstrates rendering AI agent generated UI via Google's new A2UI Compose library alongside a Koog-based agent.
Harsh Shandilya explores WorkManager 2.12.0's new ExecutionEventListener API for tracking worker execution metrics.
Place a sponsored post
We reach out to more than 80k Android developers around the world, every week, through our email newsletter and social media channels. Advertise your Android development related service or product!
alt
Libraries & Code
A zero-dependency Android library that adds shape-morphing, animated spotlight tutorials for feature discovery.
A Kotlin Multiplatform app that aggregates GitHub Trending, Hacker News, and Product Hunt with AI-powered summaries.
A real-time on-device visual editor for Jetpack Compose with zero-rebuild live UI tuning and AST-guided source splicing.
An on-device OCR SDK for Kotlin Multiplatform that extracts structured documents from PDFs and images, supporting Ukrainian.
Haze 2.0 adds a Glass effect API, a shared effects foundation for custom effects, and better blur performance.
A Kotlin Multiplatform library that monitors and inspects Ktor Client, OkHttp, and http4k network traffic.
News
Google announces the Android Auto and Automotive OS games category is now generally available for publishing.
alt
Google unveils Googlebook, Android-based laptops, and shows how adaptive layouts help apps scale to desktop.
Android Studio adds Bring Your Own Agent support, letting developers use Claude Agent, Codex, or Antigravity with IDE-native tooling.
Videos & Podcasts
Zalim Bashorov demonstrates building client-side AI-powered web apps using Kotlin/Wasm and native browser AI APIs.
Kotlin by JetBrains demonstrates the new experimental name-based destructuring feature for data classes.
Pamela Hill-Galloway shows the current state of Kotlin's Swift Export for calling shared code from Swift.
The Developers' Bakery podcast discusses Difftray, a tool for reviewing AI coding agents' changes faster than they can produce them.
Read the whole story
alvinashcraft
32 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Copilot Home, Code and Autopilot Are the Headlines. FinOps Is the Story.

1 Share

Microsoft’s September 25 announcement introduces Home, Code and Autopilot alongside a clearer commercial distinction between everyday Copilot use and usage-based agentic work.

There is plenty to get excited about. But what caught my attention is not only what Copilot can do next. It is what organizations will need to do differently.

My reading is that Microsoft is building an AI operating model—and an AI economy—inside Microsoft 365.

This is bigger than adopting another tool. It connects delegation, solution creation, autonomous work, governance and spending. And it raises a practical challenge: helping people create meaningful business value with the right AI capability, at the right cost, under the right governance model.

Let’s take a closer look.

Home and Cowork: from choosing tools to directing work

Home brings together Chat and Cowork, with Word, Excel and PowerPoint experiences integrated through Office in Copilot. Microsoft describes a future where people state what they want to accomplish and Copilot routes the work to Chat, Cowork or Code, rather than requiring them to select the mode themselves.

That is a significant direction: less focus on operating the interface, more focus on describing the outcome.

Cowork makes this shift particularly clear. Microsoft positions it around delegated, end-to-end work, including complex deliverables such as RFP responses and financial close packages.

This is not simply bigger chat.

When I delegate work, the important questions become what I want completed, which context matters, what boundaries apply and how I will judge the result. Better prompting helps, but a good prompt is not the same as a well-defined assignment.

That distinction should become part of how we teach people to work with AI.

This has already changed how I work

Cowork has already changed how I am able to work. Using Cowork and Copilot Chat on my mobile phone, I can draft, generate content and keep refining it when opening a laptop is not an option. Like preparing this blog post, planning the next work webinar, creating customer workshop materials, and the list goes on.

That might be on public transportation, sitting in a café or at home in the living room with my family. Bringing a laptop to the table with “I’ll just do some work while we’re having family night” does not really work. And yes, I should probably put the phone away as well. Touché. 🙂

The change is not just about the device. I can move work forward through conversation: describe what I need, review what comes back and steer the next iteration.

What excites me about Code and Autopilot is the possibility of extending that pattern—creating applications by chatting with Copilot, refining what I want an agent to do and reviewing its outcomes. That is the working pattern I want to build toward as these capabilities become available.

And perhaps the most useful outcome should be knowing when the work is handled—and putting the phone away.

Code: creating a solution is becoming part of everyday work

Microsoft describes Code as a way to create apps, dashboards, trackers, automations and workflows through natural language, extending solution-building beyond professional developers.

The important story is not that developers can build software. It is that business users, subject-matter experts and knowledge workers can increasingly turn their understanding of a problem into a working solution.

I see Code as extending that direction into the everyday Copilot experience. Compared with learning a visual builder or expression language, describing the desired outcome can lower the starting barrier further.

But easy to create must not become easy to abandon.

A useful team application still needs an owner, tested behavior, appropriate data permissions and a maintenance decision. A convincing first demonstration is not automatically a dependable business solution.

This is why Copilot Managed Runtime matters. Microsoft describes it as IT-governed hosting within the organization’s Microsoft 365 environment, supporting applications created through Cowork, Code and Copilot Studio; it is currently in preview.

Microsoft’s public documentation on Copilot Managed Runtime default governance settings describes controls for connectivity, sharing and runtime behavior alongside existing Power Platform governance.

For me, this is an enterprise-readiness discussion, not merely a hosting detail. Generating an application and operating it responsibly are different responsibilities.

My recommendation is to involve administrators, security teams and business owners early. Give experimentation a defined scope and a clear route from useful prototype to supported solution.

Autopilot: the work happens

Autopilot, previously known as Microsoft Scout, is Microsoft’s proactive, cloud-hosted agent for persistent work, including following up on threads, running recurring tasks and resuming projects beyond an individual interaction.

The important shift is that work can continue after the person stops interacting with it.

I find the digital-teammate framing useful, provided we do not confuse delegated execution with transferred accountability.

For persistent agentic work, I would want a business owner, a bounded objective, escalation rules, a review schedule and a clear way to stop execution.

This also connects directly to FinOps. When work continues beyond a conversation, organizations need to understand what they are funding and why.

Just because an agent keeps working does not mean the work is still needed or worth the cost.

An agent’s purpose should be reviewed alongside its quality, permissions and cost. Continuing to run is not, by itself, evidence that the work remains valuable.

The AI economy: subscription and consumption

In the Managing AI spend section of its announcement, Microsoft distinguishes between User Subscription License (USL) and Usage-Based Billing (UBB).

For everyday AI, the USL provides a fixed subscription cost covering Chat, Copilot experiences across Microsoft 365 applications, model selection and Auto model routing; Auto weighs accuracy, speed and cost when selecting a model.

Microsoft places Cowork, Code, Autopilot, long-running agentic capabilities and frontier models such as Astra and Fable under UBB. I would not frame this simply as “the interesting things cost extra.”

Someone pays for computation. If the customer is not charged separately for an operation, its cost still exists within the provider’s economics.

My view is that indefinitely expanding agentic work cannot sustainably be treated as computation without an economic consequence. Organizations should not base their strategy on that assumption. This is an economic argument, not a claim about Microsoft’s margins or unpublished pricing.

Equally, usage-based billing does not automatically mean poor value.

A demanding task can justify higher consumption if it produces a valuable, accepted result. A cheap task repeated unnecessarily can still waste money.

The useful business conversation connects the outcome, the required quality, the total cost of producing and reviewing it, and the value actually realized.

We should optimize for valuable work—not simply the lowest consumption or the most powerful model.

FinOps may be the most important announcement

The title of Microsoft’s companion announcement—New FinOps for AI capabilities: Control spend, measure value, and optimize for impact—captures the connection between financial control and business value.

I see this as a business capability, not a dashboard finance checks after IT has enabled everything.

The main Copilot announcement describes spending-policy management through APIs, credit requests routed into approval workflows and model-family controls for different user groups, including constraints on Auto’s choices.

It also describes cost-management expansion to Code and Copilot Managed Runtime, visibility into Cowork task outcomes, and users’ ability to see credit usage, remaining balances and usage history.

These are useful foundations. They are not, by themselves, proof of ROI.

A completed task is not necessarily useful work. Time saved does not automatically become financial savings. Someone still needs to establish a baseline, assess the result and decide what the organization gained.

For a pilot, I would examine accepted outputs, turnaround time, review effort, rework and consumption together. For an application, I would include maintenance and support. For autonomous work, I would also check whether the process still needs to run.

FinOps should help organizations spend confidently on valuable work—not merely spend less.

That becomes increasingly important when AI is creating applications, executing longer assignments and operating beyond individual interactions.

AI literacy needs to move beyond prompting

If an adoption program mainly teaches people to start using AI and write better prompts, I would now broaden it.

Prompting remains useful. It is simply not sufficient.

The next layer of AI literacy should include:

  • Capability selection: matching the approach to the outcome.
  • Delegation: defining objectives, boundaries and review points.
  • AI judgment: assessing evidence, quality and uncertainty.
  • Cost awareness: recognizing when additional consumption is justified.
  • Governance awareness: understanding what may be accessed, created, shared or executed.

A quick answer may not require the most advanced model. A reusable business dashboard may justify evaluating Code. A recurring process with clear boundaries may justify evaluating Autopilot when it becomes available.

These are judgment exercises, not automatic product-selection rules.

Even when Copilot handles more routing, people still need to decide whether work should be delegated and whether the result is acceptable.

My practical recommendation is to select a few meaningful outcomes, give each an owner and baseline, agree on spending and review boundaries, and scale what demonstrates value. Connect IT, finance, business owners and adoption champions rather than treating each as a separate workstream.

And do not turn cost awareness into anxiety. Give people understandable limits, room to learn and a straightforward way to request more capacity.

Rollout: keep the preview distinctions clear

Microsoft’s published rollout expectations, as of writing this article on September 27, 2026, are:

  • Home: Frontier rollout in the coming weeks.
  • Code: Frontier rollout at the end of September, with broader availability in the coming weeks.
  • Autopilot: expansion into private preview at the end of September.
  • Copilot Managed Runtime: currently in preview.
  • Plugin Registry: rolling out, with general availability across supported surfaces in the coming weeks.
  • Dynamics 365 and Power Platform grounding: public-preview rollout over the month following the announcement.

Code is also planned to enter preview for Microsoft 365 Premium and Pro subscribers later in 2026. These are consumer subscriptions, not Microsoft 365 enterprise licenses, as reflected in Microsoft’s guidance on AI credits and limits for Microsoft 365 subscriptions.

My perspective: adoption is becoming operational

I am excited about this direction. But I would not use this moment simply to add more features to a Copilot training deck. I would use it to reconsider what successful AI adoption means.

Home, Code and Autopilot are important. FinOps may prove even more important because it connects that ambition to a sustainable way of operating.

The defining challenge of the next phase is not merely getting people to use AI. It is helping them delegate responsibly, create dependable solutions and recognize which work is worth doing.

The future of work is not a maximum AI consumption nor a heavily constrained one. It is meaningful business value, created with the right AI capability, at the right cost, under the right governance model.

That is the conversation I want us to have next.



Read the whole story
alvinashcraft
32 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

How to use GitHub Copilot with WSL on Windows

1 Share
From: GitHub
Duration: 6:13
Views: 3,084

Build software on Windows using the Linux environments and tools you already rely on. In this video, see how to connect the GitHub Copilot app to Windows Subsystem for Linux (WSL) and run coding agents directly inside Ubuntu. Watch how Copilot handles parallel feature requests using worktrees, previews changes in a built-in browser, and verifies code diffs, all without leaving the app. Download the GitHub Copilot app at https://gh.io/app to get started.

— RESOURCES—

Agent automation controls in GitHub Issues: https://github.blog/changelog/2026-07-23-agent-automation-controls-in-github-issues-in-public-preview/

Creating GitHub Agentic Workflows: https://docs.github.com/en/copilot/how-tos/github-agentic-workflows/creating-github-agentic-workflows

#GitHubCopilot #Linux #WSL

— CHAPTERS —

00:00 Linux development on Windows
00:25 Connect WSL to the GitHub Copilot app
02:22 Add features with coding agents
03:35 Preview, test, and review changes

Stay up-to-date on all things GitHub by connecting with us:

YouTube: https://gh.io/subgithub
Blog: https://github.blog
X: https://twitter.com/github
LinkedIn: https://linkedin.com/company/github
Insider newsletter: https://resources.github.com/newsletter/
Instagram: https://www.instagram.com/github
TikTok: https://www.tiktok.com/@github

About GitHub
It’s where over 180 million developers create, share, and ship the best code possible. It’s a place for anyone, from anywhere, to build anything—it’s where the world builds software. https://github.com

Read the whole story
alvinashcraft
33 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

The grief, loneliness, and burnout sweeping through the tech industry right now | Molly Graham

1 Share

Molly Graham is back for round two, and this one is even more powerful. Molly has spent more than 20 years helping organizations and the humans inside them navigate growth and change. She’s held leadership roles at Google, Facebook, Quip, and the Chan Zuckerberg Initiative and is the host of TED’s WorkLife podcast (which she took over from Adam Grant). She also runs Glue Club, a leadership community for senior operators, and writes a popular newsletter called Lessons.

In our in-depth conversation, we discuss:

1. Why Molly’s famous “give away your Legos” career advice no longer holds true in an AI world

2. The grief, loneliness, and burnout sweeping through the tech industry right now

3. Why delegating to AI is fundamentally different from delegating to a human

4. The fear narrative around AI job displacement, and why it’s overblown

5. Which Legos you should never give to AI

6. What the best managers are doing right now

—

Brought to you by:

WorkOS—Make your app enterprise-ready, with SSO, SCIM, RBAC, and more

DX—Engineering intelligence for the AI era

—

Episode transcript: https://www.lennysnewsletter.com/p/the-grief-loneliness-and-burnout

Archive of all Lenny's Podcast transcripts: https://www.dropbox.com/scl/fo/yxi4s2w998p1gvtpu4193/AMdNPR8AOw0lMklwtnC0TrQ?rlkey=j06x0nipoti519e0xgm23zsn9&st=ahz0fj11&dl=0

Highlights: https://lennyspodcast.com/mostreplayedmoments

—

Where to find Molly Graham

• X: https://x.com/molly_g

• LinkedIn: https://www.linkedin.com/in/mograham

• Substack: https://mollyg.substack.com

• Website: https://glueclub.com

—

Where to find Lenny:

• Newsletter: https://www.lennysnewsletter.com

• X: https://twitter.com/lennysan

• LinkedIn: https://www.linkedin.com/in/lennyrachitsky/

—

In this episode, we cover:

(00:00) Molly Graham returns

(03:34) What is “Give away your Legos”?

(05:44) The origin story: Google, Facebook, and rapid scale

(11:27) What still holds true in an AI world

(18:13) Engineering’s identity shift: from rowing to steering

(22:09) Loneliness and the collapse of team structure

(24:06) The centaur and the reverse centaur

(25:36) The fear narrative: AI-branded layoffs and overblown doom

(29:20) Survey data: burnout is up to 55%, but half of people are thriving

(35:07) Cleaning up AI slop

(40:01) What’s actually different: delegating to AI vs. giving Legos to a human

(42:02) Why every worker is now a manager, whether they want to be or not

(49:07) Why giving things away creates space for new opportunities

(55:27) Advice for people in the age of AI

(59:38) What Legos you should never give away

(01:03:12) The human sandwich: vision at the top, AI in the middle, humans at the end

(01:05:00) Holding on to the things you love: grief, funerals, and what comes next

(01:09:39) Slow takeoff: why you’re not too late

(01:13:32) The most important skill to build right now

(01:18:58) A message for managers and leaders

(01:24:17) Key takeaways

(01:31:42) Final thoughts

—

References: https://www.lennysnewsletter.com/p/the-grief-loneliness-and-burnout

—

Production and marketing by https://penname.co/. For inquiries about sponsoring the podcast, email podcast@lennyrachitsky.com.

—

Lenny may be an investor in the companies discussed.



To hear more, visit www.lennysnewsletter.com



Download audio: https://pscrb.fm/rss/p/api.substack.com/feed/podcast/214070196/06231503de9b46eb77f3540e8404afcf.mp3
Read the whole story
alvinashcraft
33 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Running Microsoft Agent Framework Locally

1 Share

Run a full Microsoft Agent Framework (MAF) agent entirely on-device to eliminate round trips to the cloud, reduce latency, avoid egress costs, and keep sensitive data private. With Foundry Local you can run optimized local model variants (ONNX/WinML) and wire them directly into MAF.

In this guide you’ll get a working quickstart (Windows .NET 7 + WinML and a cross-platform Python example), exact commands to install and download a model, copy‑and‑paste code, hardware/driver tips, observability and troubleshooting checks, CI/CD advice, and production‑ready recommendations.

What you’ll accomplish

  • Install Foundry Local (runtime + SDK) and pick/download a local model from the Foundry catalog.
  • Create a minimal MAF agent that uses Foundry Local as the model provider.
  • Run the agent locally in C# (.NET 7 on Windows using WinML) and in Python (cross-platform).
  • Diagnose common problems (OOM, slow inference, missing model) and tune model choices.

Quick overview (high-level steps)

  1. Install Foundry Local SDK/CLI for your OS.
  2. Use the Foundry CLI or SDK to list and download a model alias.
  3. Add MAF + Foundry provider packages to your project (dotnet or pip).
  4. Configure the client (FOUNDRY_LOCAL_MODEL env var or pass at construction).
  5. Run the agent — Foundry Local will load the optimized model and MAF will call it for inference.

Prerequisites

  • OS: Windows, macOS, or Linux
  • .NET SDK: dotnet 7+ (if using C#)
  • Python: 3.8+ (if using Python)
  • Disk space: model downloads vary (see catalog); first download requires internet
  • Hardware: CPU-only works, but GPU/NPU provides much faster inference. Ensure GPU drivers & runtimes are installed (see GPU section below).
  • Command-line familiarity, ability to set environment variables

Quickstart (10–15 commands)
This is a condensed path from zero to a running agent (Windows example for .NET then Python). If you prefer a single OS, follow that section below in full.

Windows (.NET 7 + WinML quick path)

  1. Install dotnet 7 SDK (if not installed).
  2. Create a new console app and add packages:
    • dotnet new console -n FoundryAgentDemo
    • cd FoundryAgentDemo
    • dotnet add package Microsoft.Agents.AI
    • dotnet add package Microsoft.Agents.AI.Foundry –prerelease
    • dotnet add package Microsoft.AI.Foundry.Local.WinML
  3. Set model alias (PowerShell):
    • $env:FOUNDRY_LOCAL_MODEL = “phi-4-mini”
  4. Build & run:
    • dotnet build
    • dotnet run
      (On first run Foundry Local will download the chosen model. Expect time for the initial download.)

Cross-platform Python quick path

  1. Create virtualenv and install (adjust package names if necessary — see details below):
    • python -m venv venv && source venv/bin/activate
    • pip install agent-framework foundry-local-sdk
  2. Set model alias (bash):
    • export FOUNDRY_LOCAL_MODEL=”phi-4-mini”
  3. Run main.py (see Python example below)

Important note: package and CLI names evolve. If a command fails, consult the Foundry Local docs and the official sample repos (foundry-samples and agent-framework) which contain tested examples and exact package/namespace names.

Part A — Install Foundry Local and download a model (details)

  1. Install options
    Foundry Local can be consumed via:
  • NuGet packages for .NET (recommended for C# apps)
  • PyPI packages (for Python)
  • npm packages (for JS)
  • Native installers or a platform-specific CLI where available

Where to get the right installer/packages

  • Microsoft Learn Foundry Local get-started (follow the platform-specific instructions)
  • Official GitHub samples: microsoft-foundry/foundry-samples and microsoft/agent-framework
  1. Example .NET installation (project-level)
    You don’t usually “install” the runtime globally for .NET; add the Foundry Local package(s) to your project:
  • dotnet add package Microsoft.AI.Foundry.Local (Linux/macOS/CPU/GPU)
  • On Windows to use WinML/DirectML acceleration:
    • dotnet add package Microsoft.AI.Foundry.Local.WinML
  1. Example Python installation
  • pip install foundry-local-sdk
  • pip install agent-framework
    (Exact pip package names may vary across releases; the sample repo shows exact names. If pip cannot find the package, clone/from-source in the foundry-samples repo.)
  1. The Foundry CLI (catalog management)
    Foundry exposes a catalog of local model aliases and (often) a CLI for listing/downloading models. CLI names vary by release (examples you may see in docs: foundryctl, foundry). Typical CLI flow:
  • foundryctl catalog list
  • foundryctl catalog inspect phi-4-mini
  • foundryctl model download phi-4-mini –path /path/to/foundry/models

If your platform’s docs show a different CLI name, use that. The initial model download can be large; download location and cache are configurable in the runtime docs.

  1. Choosing a model
  • Check alias metadata for VRAM/disk estimates and license.
  • If you have limited VRAM, pick smaller or quantized variants.
  • Common small aliases: phi-4-mini, phi-2-mini, qwen-7b-quant (names change — check the catalog).

Part B — Copy‑paste ready C# example (Windows .NET 7 + WinML)

Important: package names and API surfaces can change between SDK releases. The commands below use package names commonly seen in the Microsoft samples. If NuGet packages change, run the dotnet add package commands without versions to get the latest packages, and refer to the foundry-samples repo for exact Program.cs if you encounter API differences.

  1. Project setup (commands)
  • dotnet new console -n FoundryAgentWinML
  • cd FoundryAgentWinML
  • dotnet add package Microsoft.Agents.AI
  • dotnet add package Microsoft.Agents.AI.Foundry –prerelease
  • dotnet add package Microsoft.AI.Foundry.Local.WinML
  • (optional) dotnet add package Microsoft.Extensions.Hosting
  1. Example Program.cs (copy into Program.cs)
    (This example shows the pattern: create Foundry local client, wrap in MAF chat agent, and run a prompt. Adjust namespaces if your SDK version uses slightly different names.)

using System;
using System.Threading.Tasks;
using Microsoft.Agents.AI;
using Microsoft.Agents.AI.Foundry;

class Program
{
static async Task Main(string[] args)
{
// Read model alias from environment or use a default
var modelAlias = Environment.GetEnvironmentVariable(“FOUNDRY_LOCAL_MODEL”) ?? “phi-4-mini”;

    Console.WriteLine($"Using Foundry local model: {modelAlias}");

// Create a FoundryLocalClient (provider glue). Exact constructor may vary by SDK version.
var foundryClient = new FoundryLocalClient(model: modelAlias);

// Create agent wrapper using a ChatClientAgent helper (example API)
var agentOptions = new ChatClientAgentOptions
{
SystemPrompt = "You are a helpful assistant that summarizes text concisely."
};

var agent = ChatClientAgent.FromChatClient(foundryClient, agentOptions);

var input = "Summarize the plan for the project in one sentence.";
Console.WriteLine($"Sending prompt: {input}");

var result = await agent.RunAsync(input);
Console.WriteLine("Agent response:");
Console.WriteLine(result);
}
}

  1. Build and run
  • Open PowerShell and set the model alias:
    • $env:FOUNDRY_LOCAL_MODEL = “phi-4-mini”
  • dotnet build
  • dotnet run

Notes

  • On first run, Foundry Local will download and prepare model artifacts. Check the console logs for model download progress.
  • If you want to pass the model explicitly instead of using an env var, call the FoundryLocalClient constructor with the desired alias (see SDK docs).

Part C — Copy‑paste ready Python example (cross-platform)

  1. Create virtual environment and install (adjust if package names differ)
  • python -m venv venv
  • source venv/bin/activate (macOS/Linux)
  • venv\Scripts\Activate.ps1 (Windows PowerShell)
  • pip install agent-framework foundry-local-sdk

If PyPI names differ or packages aren’t present, clone the sample repo and install from source.

  1. requirements.txt (example)
    agent-framework
    foundry-local-sdk
  2. main.py (copy into your project)
    import os
    import asyncio

from agent_framework import Agent
from agent_framework.foundry import FoundryLocalClient

async def main():
model_alias = os.environ.get(“FOUNDRY_LOCAL_MODEL”, “phi-4-mini”)
print(f”Using Foundry local model: {model_alias}”)

# Create Foundry client (constructor may vary)
client = FoundryLocalClient(model=model_alias)

# Create a simple agent wrapper — adjust arguments to match the library you installed
agent = Agent(chat_client=client, instructions="You are a brief assistant that summarizes text.")

response = await agent.run("Explain Docker in one sentence.")
print("Agent response:", response)

if name == “main“:
asyncio.run(main())

  1. Run
  • export FOUNDRY_LOCAL_MODEL=”phi-4-mini” (macOS/Linux)
  • $env:FOUNDRY_LOCAL_MODEL = “phi-4-mini” (Windows PowerShell)
  • python main.py

Notes

  • Use the sample repo’s Python examples to get exact import paths if the package you installed exposes different modules.
  • On first run the Foundry runtime will download the model. Monitor logs.

Part D — Configure Foundry Local without env var (programmatic control)
Instead of environment variables, pass the model alias when you create the client:

  • C#:
    var foundryClient = new FoundryLocalClient(model: “qwen2.5-1.5b-instruct”);
  • Python:
    client = FoundryLocalClient(model=”qwen2.5-1.5b-instruct”)

This is useful for runtime selection, multi-agent scenarios, or when you want config inside your app.

Hardware, drivers and model sizing (practical guidance)

  • NVIDIA GPUs: install NVIDIA drivers + CUDA toolkit compatible with the Foundry runtime. Many ONNX/WinML optimizations target CUDA 11.x; check Foundry docs for the exact CUDA/cuDNN versions required for a given release. Verify installation with:
    • nvidia-smi
  • AMD GPUs: ROCm support varies by OS and GPU generation. Check ROCm compatibility guide.
  • Windows DirectML / WinML: WinML provides DirectML acceleration; keep Windows and GPU drivers up-to-date. For Intel GPUs, install the latest Graphics drivers.
  • VRAM and disk: model sizes and VRAM requirements are model-specific. Smaller models and quantized variants (4-bit/8-bit) are recommended for constrained devices. Use the Foundry catalog to inspect model metadata (size and suggested device types).
  • If GPU memory is insufficient, pick (a) a smaller model, (b) a quantized variant, or (c) fall back to CPU execution (slower).

What to look for in the Foundry catalog

  • Alias name
  • Disk size and estimated VRAM
  • Quantization variants offered (4-bit/8-bit)
  • License and usage restrictions

Observability and logging

  • Enable MAF OpenTelemetry integration to capture agent spans (system prompt, tool calls, model inference).
  • Foundry Local logs (model download, provider selection, errors) appear in the runtime output. If you start Foundry as a background service, logs typically go to console files or a configured log directory — check the Foundry docs for the exact log file location on your OS.
  • Example telemetry flow (conceptual):
    • Start tracing in your app (OpenTelemetry SDK).
    • Agent spans: pre-processing (prompt assembly), model inference (Foundry span), post-processing (tool calls).
    • Use traces to identify slow stages (model load vs. token-by-token decoding).

Troubleshooting checklist (concrete remedies)

  1. Model not found / invalid alias
  • Confirm alias with the Foundry catalog CLI:
    • foundryctl catalog list
    • foundryctl catalog inspect
  • Ensure FOUNDRY_LOCAL_MODEL is set correctly.
  1. First run stalls on download
  • Check network connectivity and disk space.
  • Watch Foundry log lines showing download progress.
  • Consider pre-downloading models on CI machines to avoid repeated downloads.
  1. Slow inference
  • Confirm GPU provider selection: check Foundry logs for “selected execution provider” lines.
  • Ensure GPU drivers and runtimes are properly installed (nvidia-smi, rocminfo).
  • Use a smaller or quantized model if GPU not available.
  1. Out-of-memory / VRAM errors
  • Use smaller model alias or quantized variant.
  • Use CPU execution for lower memory devices.
  • If running multiple models concurrently, reduce concurrency or run only one model at a time.
  1. API/namespace mismatches
  • If compilation fails due to missing types/namespaces, check the sample repo for the exact package and API names for your SDK version. The foundry-samples and agent-framework samples are the definitive, copy‑paste-ready references.
  1. Tooling and hosted features unavailable
  • Foundry Local is a local chat client. Some hosted Foundry tools (hosted web search, hosted code interpreter) are not available locally. Build local replacements or use Foundry hosted services when you need managed tools.

Forensics & logs: where to look

  • Application logs (your process) — model client creation and errors.
  • Foundry runtime logs — download progress, execution provider selection.
  • OS and GPU driver logs (nvidia-smi, dxdiag on Windows).
  • CI logs for model download step (cache model artifacts to speed repeated runs).

CI/CD and production considerations

  • Avoid repeated large downloads in ephemeral CI runners by caching the model artifacts. Use a build cache or artifact store with pre-downloaded model files.
  • For smoke tests, use small local models or mocked chat clients to validate agent logic without heavy downloads.
  • Memory budgeting: if multiple agents or other services use the same GPU, coordinate startup to avoid OOM.
  • Hybrid deployment: run smaller on-device models for latency/privacy, and use hosted Foundry for heavy tooling or when you need managed, up-to-date models and scalable inference.

Security and licensing checklist

  • Check model license in the Foundry catalog — confirm commercial usage terms if applicable.
  • Store no secrets in code. Foundry Local does not require cloud API keys for local inference, but MAF or other integrations might — use secure stores.
  • Limit permissions on model cache directories and follow OS best practices for process isolation.

Performance tuning tips

  • Batch inference when possible (if the agent supports multiple concurrent requests).
  • Use quantized model variants for lower memory and faster CPU inference; be conscious of potential accuracy trade-offs.
  • Prefer GPU acceleration when available — it typically reduces latency per request significantly.
  • Warm-up the model after download by running a small inference to optimize runtime caches.

Appendix: Useful command snippets (summary)

Set env var

  • Windows PowerShell:
    • $env:FOUNDRY_LOCAL_MODEL = “phi-4-mini”
  • macOS / Linux:
    • export FOUNDRY_LOCAL_MODEL=”phi-4-mini”

Dotnet project (example)

  • dotnet new console -n FoundryAgentDemo
  • dotnet add package Microsoft.Agents.AI
  • dotnet add package Microsoft.Agents.AI.Foundry –prerelease
  • dotnet add package Microsoft.AI.Foundry.Local
  • dotnet run

Python virtualenv

  • python -m venv venv && source venv/bin/activate
  • pip install -r requirements.txt
  • python main.py

Foundry CLI (example patterns — CLI name may vary)

  • foundryctl catalog list
  • foundryctl catalog inspect phi-4-mini
  • foundryctl model download phi-4-mini –path /path/to/models

Where to get canonical examples

  • Foundry-samples on GitHub: microsoft-foundry/foundry-samples (C# and platform-specific samples)
  • Agent Framework repo: microsoft/agent-framework (sample projects for C# and Python)
  • Microsoft Learn Foundry Local docs: (Foundry Local get-started & SDK reference)
    If any command or API fails, consult those repos for exact, copy-paste-ready Program.cs and main.py.


Foundry Local + MAF gives you an approachable path to local agents that respect privacy and latency constraints. The main friction is model selection and the first-time download/optimization. Start with the official sample repos (foundry-samples and agent-framework) for exact API versions, try a small model (phi-4-mini) for initial experiments, and iterate to larger/quantized models as your hardware allows

Note: The initial draft of this post was written by BlogWriter and then edited by Jesse Liberty
Illustrations by Copilot. Caution: LLMs make mistakes; this post is offered as is.

Read the whole story
alvinashcraft
33 minutes ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories