Executive Summary, Day One
📠Inbox 5:05AM
To: Ternus, John <jternus@apple.com>
Reply-to: CEO-transition-team@apple.com
Welcome to your first day as CEO! Hopefully you're finding your new office spacious and comfortable. If you have trouble with any of the doors just ping us - being a Jony Ive design, you won't find anything as unsightly as a door handle here!
Today's memo includes a summary of the business, broken out by division. The day's schedule is also attached; please note that the Apple Intelligence team needed to delay the new Siri discussion until a later date, but they promise they're "for real" this time.
Hardware
Everything …
Today, Microsoft published its 2026 Responsible AI Transparency Report. The report highlights the progress we’ve made in building and deploying AI responsibly, supporting our customers, and strengthening our responsible AI governance, tools, and practices. You can explore the report in its entirety here.
AI is moving fast, and so are societal expectations. The boundaries of what people can accomplish with AI are expanding, and communities are asking more questions about how AI systems are designed, built, and used. As AI becomes more integral to how we live and work, confidence that AI systems are operating reliably and securely is becoming an essential prerequisite to their broad and beneficial adoption.
At Microsoft, we have been building a responsible AI program for nearly a decade, rooted in two core beliefs: that trust is foundational to realizing the benefits of AI and that the empowerment of people and organizations must remain at the center of our strategy. As capabilities advance and adoption accelerates, that experience is helping us meet this moment and adapt for what comes next.
Our third annual Responsible AI Transparency Report shares how our program is evolving and the priorities that continue to shape our work. Over the last year, investing in three specific areas has enabled us to embed trust more deeply and at greater scale: adaptive governance and technical risk management, practical tools and capabilities, and shared practices and strong partnerships. These investments cut across five trends shaping the AI landscape, including the rapid expansion of agentic AI.
Taken together, these trends and investments underscore our view that model capability alone will not determine the impact of AI. That will depend on organizations that develop and deploy AI technologies that deliver real value—and govern them with the rigor and adaptability needed to earn and sustain trust.
As the frontiers of AI advance, we are making our governance more adaptive and more tightly integrated with engineering workflows. In practice, this means that we have updated our policies to better match the AI tech stack and AI value chain, evolved our risk management practices to address emerging AI capabilities and risks, and strengthened the readiness of our responsible AI community: the people who operationalize our program at enterprise scale.
This year, we re-engineered our Responsible AI Standard to make it more adaptive to evolving technical realities, uses, risks, and regulatory requirements. The new Standard is structured by reference to different components of the tech stack—models, platform services, and applications—and the role that Microsoft plays in developing or deploying those components. It combines core requirements that always apply with more targeted, scenario-specific requirements that can evolve as capabilities and risks change. For example, we apply some of our most rigorous risk management measures to AI systems with the most significant cyber capabilities, helping ensure that advances in AI favor the defenders responsible for securing critical digital infrastructure.
We are also evolving our technical risk management practices. Increasingly capable systems can retain memory, use tools, access data, and take actions on behalf of users. Governing these systems requires us to think beyond the behavior of an individual model or application to interactions among models, agents, applications, tools, data, and people. Our work increasingly focuses on controls such as agent identities, tool permissions, and monitoring of actions.
And governance only works when people can put it into practice. We have continued to build responsible AI capabilities across Microsoft, equipping thousands of engineers and product managers with training on topics such as agentic AI threat modeling and prompt injection defenses.
Together, these investments are helping us move toward a more continuous, lifecycle-based approach to AI governance. With agentic AI, risks can evolve as systems interact with their environments, users, and other systems. Our governance needs to evolve with these agentic capabilities—and incorporate what we learn from their testing and deployment.
Effective governance depends on tools that help translate policy goals into action. As developers and organizations navigate a more complex technical and regulatory environment, they need practical ways to identify risks, evaluate systems, establish controls, and monitor how AI behaves in the real world.
We are applying what we learn from governing AI at Microsoft into tools, capabilities, and resources that help developers and organizations beyond Microsoft do just that—whether they build on our platforms or leverage open-source projects.
We have expanded tools to evaluate AI systems across the lifecycle. A new AI Red Teaming Agent helps accelerate the identification and evaluation of risks. Agent evaluators help developers measure the quality, safety, and performance of agentic applications. RAMPART turns red team findings into repeatable tests, enabling more continuous coverage as systems change.
We are also building greater visibility and control into agentic systems. With ASSERT and Agent Control Specification, developers can evaluate agents against their policies, place runtime controls at critical points in an agent’s workflow, and monitor behavior.
These tools and capabilities reflect a shift: as systems become more dynamic, governance needs to become more operational. Organizations need to be able to see what their systems are doing, test how they behave, and intervene when necessary—not just assess them before deployment.
Organizations also need confidence—and increasingly need to demonstrate—that responsible AI practices are being implemented consistently. Microsoft is one of the few companies certified against ISO 42001 across a broad portfolio, including Microsoft 365 Copilot, Foundry, and GitHub Copilot. Over the last year, we have simplified and strengthened our internal processes that support that certification.
Ultimately, responsible AI governance is a shared responsibility across the AI value chain. Our goal is to help make the practices and capabilities needed to meet that responsibility more accessible, practical, and scalable.
The challenges of governing AI are bigger than any one company, and increasingly interconnected AI systems make collaboration even more essential.
As AI adoption expands across borders and sectors, we need shared expectations for how systems are evaluated, monitored, and governed, as well as interoperable standards that enable visibility into interactions across tools, data, and systems. We also need to keep advancing the underlying science and technical practices so that we can benefit from rigorous, applied insights into what effective governance looks like and where the remaining gaps are.
That starts with research. Over the past year, we advanced our work with the US Center for AI Standards and Innovation and AI Safety and Security Institutes in Australia, Singapore, and the UK to strengthen the science and practice of AI evaluation. We also launched an External Red Team Alliance with 18 universities across six continents to expand understanding of priority risks.
Common technical practices and standards are critical. Through the Frontier Model Forum, OpenTelemetry, and the Appia Foundation, we are helping develop approaches spanning frontier cyber benchmarks, end-to-end observability for increasingly agentic systems, and AI assurance across supply chains and sectors. We are also contributing to efforts that make transparency reporting more interoperable across organizations and jurisdictions, including through an OECD-led informal task force that developed the Hiroshima AI Process Reporting Framework version 2.0.
We also need shared ways to measure progress. We cannot meaningfully assess progress if every organization measures AI risks differently. Through our work with MLCommons, we are helping expand AILuminate into a broader suite of reliability benchmarks, creating common approaches for evaluating areas such as jailbreak resilience, multilingual performance, and psychosocial risk in conversational AI.
Shared learning, shared practices and standards, and shared measurement can help the entire ecosystem develop while raising shared expectations for trust.
Our experience over the past year has reinforced that responsible AI cannot be static. It has to be embedded in development processes, supported by practical tools, and continually informed by what we learn. That is why our responsible AI investments extend from the systems we build, to the tools we provide our customers, to the research, practices, and measurement approaches we help develop with the broader ecosystem.
Our 2026 Responsible AI Transparency Report explores this work in more depth—from how we re-engineered our Responsible AI Standard to how we are strengthening governance for agentic AI, advancing evaluation, and addressing AI misuse. We invite you to explore the report to see what we have learned, what we have changed, and how we are putting our priorities into practice.
As AI becomes more powerful and more present in people’s lives, our commitment is to keep listening and learning, to keep strengthening our safeguards, and to keep putting the empowerment of people and organizations at the center of our strategy.
The post Responsible AI in 2026: How we are adapting for what’s ahead appeared first on Microsoft On the Issues.
Coauthored with Claude
Midway through each month, I think “The next Trends is going to be small. Not much is happening.” This is the first time that I’ve been right. Was everyone on vacation in August? Am I becoming jaded? There were many model releases, though few of them seemed significant. Then again, it may be time to get over the one-upmanship by the frontier vendors and spend more time thinking about the myriad small and open-weight models. Every month, the best laptop-scale models (30B and smaller) seem closer to the leading frontier models. And every month, we’re seeing organizations realize that paying premium per-token prices for the latest frontier models gives at best a small advantage over the best open-weight models.
Capability and model size are decoupling. Several models here run comfortably on a laptop or a single accelerator while claiming performance close to much larger frontier systems. While it can be hard to work with a smaller model without thinking that you’re choosing “second best,” the biggest model isn’t always the right choice. Major releases aside, the most important news from August might be Anthropic’s deployment of watermarks for text. If the watermarking scheme works, it will be possible to tell which parts of an article like this were written by AI.
Features that we associate with agents or harnesses, such as the ability to spawn subagents and delegate tasks to less-expensive models, are continuing to find their way into the models themselves. There’s also a countertrend: Individuals and organizations are building their own agents that are closely integrated into their working environment. Are we headed for walled gardens controlled by the leading providers? Or will a thousand flowers bloom, each reflecting an idiosyncratic way of working with AI? Don’t avoid tools from the major AI labs, like Claude Code and Codex, but don’t lock yourself into thinking that they’re the only option.
Optimizing AI usage has become its own discipline, sometimes called “tokenomics.” Tokenomics can’t be separated from safety, which has also been much in the news. Disposable containers built for agents, GPU scheduling that treats accelerators as a heterogeneous pool, and infrastructure providers publishing how they actually serve open models at scale all match workloads to hardware without waste or risk. AI performance isn’t just about models; it’s about infrastructure. Understanding how the model is run will prove more important than the model’s specs and benchmarks.
Security work is inseparable from AI development, not a layer added afterward—but security professionals have been saying that about traditional software for years. Artificial intelligence is spawning new attacks as well as new defenses. While it’s always fascinating to look at new attacks, the most significant shift is in defense: rethinking security in terms of actions and resources rather than user identities, a change we’ve also covered on the Radar blog.
How do people use AI? Does AI use lead to greater productivity? We know surprisingly little about either question. We’re still learning how to use AI effectively; the best metric isn’t a simple measure of productivity but whether you can do things you couldn’t do before.
There’s now a specialized version of ChatGPT for teens; a site that serves different content to scrapers and humans; and an AI-generated animation of the start of The Lord of the Rings. The web is proving that it can adapt to anything that’s thrown at it. It’s where we learn and play, and AI isn’t changing that.
In my previous article on Dev.to, I showed how to build expressive text-to-speech using Gemini and Firebase Cloud Functions. While that setup worked beautifully, it required building and deploying custom backend endpoints. Firebase AI Logic enables running Gemini TTS client-side with production security, eliminating custom backend code and automating deployment via Git integration.
Project technical stack:
The public Google Gemini Developer API is restricted in my region (Hong Kong). However, the Agent Platform Gemini API (Google Cloud) offers enterprise access that works reliably here, so I chose the Agent Platform Gemini API for this Firebase AI Logic demo.
npm install -g firebase-tools
Install or update firebase-tools globally using npm.
firebase logout
firebase login
Log out and re-authenticate with Firebase.
npm i --save-exact firebase
npm i --save-exact --save-dev firebase-tools serve
Install the dependencies to call the Firebase AI Logic API, use the Firebase CLI to generate files, and serve the production build.
firebase init
Execute firebase init and follow the prompts to set up the Firebase AI Logic, Emulators, App Hosting, and Remote Config.
If you have an existing project or multiple projects, you can specify the project ID on the command line.
firebase init --project <PROJECT_ID>
After completing the setup steps, the Firebase tools generate the configuration files such as .firebaserc and firebase.json. You can view the .firebaserc and firebase.json in the GitHub repo.
We asked antigravity-cli (a terminal-first AI coding agent released by Google) and the Gemini Flash model to create two Node.js scripts to do the following:
public/remote-config-defaults.json. You can read the complete listing of the script.public/firebase.config.json. If we hardcode the public keys and push them to GitHub, GitHub triggers a false alarm. You can read the complete listing of the script and the environment variable template.
# generated firebase configuration
firebase.config.json
Add firebase.config.json to .gitignore to prevent accidental commits.
During build time, Angular bundles both JSON files into the dist directory to provide initial configuration values.
Users submit text to Firebase AI Logic to synthesize speech. Firebase generates the complete L16 audio payload and returns it to the client. Because the HTML audio element does not support the L16 format, the application converts the audio to a WAV Blob before binding the Blob URL to the element source.
The second flow streams the audio and sends the L16 chunks to the Angular application. The audio player creates an AudioBufferSourceNode to play the chunk data and cleans up resources to prevent memory leaks.
While the full codebase is available in the ng-firebase-tts repository, the application relies on Firebase Remote Config to manage configuration, App Check to prevent abuse, and App Hosting to deploy the Angular application.
The following sections illustrate how to initialize the Firebase app and App Check, and how to activate Remote Config values.
The public/firebase.config.json file contains the public Firebase API key, the sensitive reCAPTCHA Enterprise key, and the App Check debug token to bypass device attestation in local development. These values are critical in Firebase App and App Check initialization.
@Service()
export class ConfigService {
#app: FirebaseApp | undefined = undefined;
#remoteConfig: RemoteConfig | undefined = undefined;
/*... getter methods are omitted... */
get appConfig(): AppRemoteConfig {
return this.#appConfig;
}
get aiBackend(): AI {
return this.#aiBackend;
}
async initialize(): Promise<void> {
this.#app = initializeApp(firebaseConfig.app);
(globalThis as any).FIREBASE_APPCHECK_DEBUG_TOKEN = firebaseConfig.appCheckDebugToken || true;
initializeAppCheck(this.#app, {
provider: new ReCaptchaEnterpriseProvider(firebaseConfig.recaptchaEnterpriseKey),
isTokenAutoRefreshEnabled: true,
});
this.#remoteConfig = getRemoteConfig(this.#app);
this.#remoteConfig.defaultConfig = remoteConfigDefaults;
await fetchAndActivate(this.#remoteConfig);
this.#appConfig = {
vertexAILocation: getValue(this.#remoteConfig, 'vertexAILocation').asString(),
useLimitedUseAppCheckTokens: getValue(
this.#remoteConfig,
'useLimitedUseAppCheckTokens',
).asBoolean(),
geminiTTSModelName: getValue(this.#remoteConfig, 'geminiTTSModelName').asString(),
};
this.#aiBackend = getAI(this.#app, {
backend: new AgentPlatformBackend(this.#appConfig.vertexAILocation),
useLimitedUseAppCheckTokens: this.#appConfig.useLimitedUseAppCheckTokens,
});
}
}
The initialize method initializes Firebase App, configures App Check, sets up Firebase AI, and assigns Remote Config values to appConfig.
The appConfig object holds the location, the Gemini TTS model name, and the limited-use App Check token flag.
Configure the TTS model name and the limited-use App Check token parameters in Firebase Remote Config. Both parameters have conditional values. When the Firebase web application is firebase-ai-logic-tts, the TTS model is gemini-3.1-flash-tts-preview, and the limited-use App Check token flag is set to true.
export const AI_BACKEND = new InjectionToken<AI>('AI_BACKEND');
export function provideFirebase() {
return makeEnvironmentProviders([
{
provide: AI_BACKEND,
useFactory: () => inject(ConfigService).aiBackend,
},
]);
}
The AI_BACKEND injection token provides a factory function to return the Firebase AI from ConfigService.
export const appConfig: ApplicationConfig = {
providers: [
... other providers ...
provideAppInitializer(async () => await inject(ConfigService).initialize()),
provideFirebase(),
],
};
provideAppInitializer and provideFirebase initialize Firebase and configure Firebase AI during application startup.
The private createModel method calls getGenerativeModel to create a generative model configured with speechConfig. The subsequent workflows use this model to convert L16 audio to WAV or stream L16 playback directly.
@Service()
export class TextToSpeechService {
readonly #configService = inject(ConfigService);
readonly #modelName = this.#configService.appConfig.geminiTTSModelName;
private createModel(voiceName: string) {
return getGenerativeModel(this.#aiBackend, {
model: this.#modelName,
generationConfig: {
responseModalities: [ResponseModality.AUDIO],
speechConfig: {
voiceConfig: {
prebuiltVoiceConfig: { voiceName },
},
languageCode: 'en-US',
},
},
});
}
}
async synthesize(text: string, voiceName: string): Promise<Blob> {
const model = this.createModel(voiceName);
const result = await model.generateContent([text]);
const chunk = this.extractValidChunkData(result.response);
const { data, mimeType } = chunk;
return convertToWav(decodeBase64(data), mimeType);
}
The extractValidChunkData method extracts binary audio data and the MIME type from the response payload.
The synthesize method retrieves the entire audio payload in L16 format. However, the HTML audio element does not support L16, so the application must convert the data to WAV format before assigning the Blob URL to the element source.
See the conversion to WAV code for the full implementation.
This implementation remains straightforward because it avoids managing response streams and incremental audio chunks. However, users suffer from latency when the text is long and it produces a long audio stream. To eliminate playback latency, the next section explores streaming chunks directly via the Web Audio API AudioContext.
async *synthesizeStream(text: string, voiceName: string): AsyncGenerator<RawAudioBinary | undefined> {
const model = this.createModel(voiceName);
let firstMimeType = '';
let sampleRate = DEFAULT_SAMPLE_RATE;
const responseStream = await model.generateContentStream([text]);
for await (const chunk of responseStream.stream) {
const chunkData = this.extractValidChunkData(chunk);
if (chunkData) {
const { data, mimeType } = chunkData;
const decodedData = decodeBase64(data);
if (!firstMimeType && mimeType) {
firstMimeType = mimeType;
sampleRate = parseMimeType(firstMimeType).sampleRate;
}
yield { decodedData, sampleRate };
}
}
yield undefined;
}
The synthesizeStream method returns an asynchronous generator that yields the raw binary data and the sample rate. The MIME type is audio/l16; rate=24000; channels=1, so parseMimeType extracts the sample rate for the audio player's AudioContext.
While AudioPlayerService is beyond the scope of this article, the source code demonstrates how to play back audio without using an HTML audio element.
Next, we will build a reactive user interface in Angular that renders an HTML audio element to play audio.
The TextToSpeechComponent delegates audio management to the view service to maintain separation of concerns.
The TextToSpeechComponent displays three buttons to generate speech from text across three different scenarios:
@Component({
selector: 'app-text-to-speech',
templateUrl: './text-to-speech.component.html',
styleUrl: './text-to-speech.component.css',
imports: [SpinnerIconComponent, NgTemplateOutlet],
providers: [TextToSpeechViewService],
})
export class TextToSpeechComponent {
private readonly speechService = inject(TextToSpeechViewService);
interestingFact = input<string | undefined>(undefined);
audioPrompt = input.required<string>();
voice = input.required<string>();
async generateSpeech(mode: GenerateSpeechMode) {
const fact = this.interestingFact();
await this.speechService.generateSpeech(mode, {
prompt: this.audioPrompt(),
voice: this.voice(),
fact: this.interestingFact(),
});
}
}
The TextToSpeechViewService encapsulates TextToSpeechService and AudioPlayerService to coordinate speech synthesis and audio playback.
@Injectable()
export class TextToSpeechViewService {
private readonly speechService = inject(TextToSpeechService);
private readonly audioPlayerService = inject(AudioPlayerService);
#audioUrl = signal<string | undefined>(undefined);
audioUrl = this.#audioUrl.asReadonly();
private processStreamChunk(isInitialized: boolean, playbackRate: number, chunk: RawAudioBinary) {
if (!isInitialized) {
this.audioPlayerService.initialize(chunk.sampleRate, playbackRate);
isInitialized = true;
}
this.audioPlayerService.processChunk(chunk.decodedData);
return isInitialized;
}
private async handleSync(promptArgs: FactConfig) {
const blob = await this.speechService.synthesize(promptArgs.prompt, promptArgs.voice);
this.setAudioUrl(blob);
}
private async handleStream(promptArgs: FactConfig) {
let isInitialized = false;
const { prompt, voice } = promptArgs;
for await (const chunk of this.speechService.synthesizeStream(prompt, voice)) {
isInitialized = this.processStreamChunk(isInitialized, 1, chunk);
}
}
private setAudioUrl(finalBlob: Blob | undefined) {
if (finalBlob) {
const createdUrl = URL.createObjectURL(finalBlob);
this.#audioUrl.set(createdUrl);
return createdUrl;
}
return undefined;
}
async generateSpeech(mode: GenerateSpeechMode, promptArgs: FactConfig) {
revokeBlobURL(this.#audioUrl());
this.#audioUrl.set(undefined);
switch (mode) {
case 'sync':
await this.handleSync(promptArgs);
break;
case 'web_audio_api':
await this.handleStream(promptArgs);
break;
}
}
}
handleSync invokes the Firebase SDK to synthesize audio using the Gemini TTS model, generates a Blob URL, updates #audioUrl, and renders the HTML audio element.
handleStream calls the Firebase SDK to use the Gemini TTS model to stream the audio instead. The audio context receives the chunk data, plays it immediately, and avoids rendering the HTML audio element.
The integration of text-to-speech with Firebase AI Logic empowers Angular applications for real-time audio generation.
The Angular application handles text-to-speech entirely on the client. Pushing changes to Git triggers an automatic deployment to Firebase App Hosting.
Try cloning the GitHub repository, uploading an image to generate an obscure fact, and using the Gemini 3.1 Flash TTS preview model to speak it with the specified scene, emotion, and pace.