President Donald Trump has a knack for turning words against his enemies. His first successful presidential run was built on monikers like "Little Marco" and "Crooked Hillary"; he changed "fake news" from a phrase describing scammy media outlets to a derogatory term for the press at large. Over the past few weeks, he's clearly decided he can work the same magic to promote artificial intelligence - branding a technology he wants to accelerate "super", while turning "artificial" into his latest go-to pejorative and declaring resisters "THE ENEMY." Some of the biggest names in AI are going along with him. But he's picked a tough linguistic batt …
GitHub COO Kyle Daigle, who now also leads developer marketing for all of Microsoft, speaks at GitHub Universe last year. (Microsoft Photo)
GeekWire is profiling over the next few weeks some of the people and teams that are shaping the evolution of Microsoft in what we’re calling its “Microsoft 2.5” era.
A delicate dance: In August 2025, Microsoft moved to integrate GitHub into the mothership, under its CoreAI division, after having allowed it to run as a quasi-independent entity since acquiring it in 2018. At the same time, GitHub CEO Thomas Dohmke announced he would leave the company at the end of the year.
A number of developers feared the worst: GitHub would become just another arm of Microsoft and lose the community-mindedness that had made it attractive to developers of all stripes, and especially open-source ones.
GitHub Chief Operating Officer Kyle Daigle — a 13-plus-year GitHub veteran — had been one of the main champions of the need to “keep GitHub GitHub.” He maintained that stance even after he added the “Chief Marketing Officer of Developer” title for all of Microsoft to his COO role in January 2026.
So, how has Microsoft done this past year with managing GitHub? In an interview with GeekWire, Daigle acknowledged that GitHub’s relationship with Microsoft has changed since the end of the standalone CEO era. But the goal is no longer just “don’t break GitHub.” It’s to export GitHub’s community, developer-first ethos and learnings across Microsoft.
GitHub has influence beyond just the GitHub product set now, Daigle said. And the fact that Microsoft opted to consolidate its developer-facing messaging and community outreach under Daigle, someone who came from GitHub, not from Microsoft’s traditional DevDiv organization, gives weight to his claim.
“Bringing teams together — engineering teams and marketing teams and everyone that wasn’t talking to each other before — has been a really big part of the work,” Daigle said. “I see all these opportunities for GitHub to directly help and impact the overall mission of Microsoft versus keeping it cloistered in a way that isn’t helping either GitHub’s mission or Microsoft’s mission.”
These days, developer teams across Microsoft and GitHub are sharing foundations and the GitHub Copilot software development kit across organizations “in a way that would have seemed improbable before,” he said.
As Microsoft historians know, “Microsoft’s challenge isn’t creating products. It’s connecting them,” Daigle said.
Growing pains: At the same time, the rise of AI agents has strained GitHub’s infrastructure. Over the past few months, GitHub has experienced some significant outages, including one in August that lasted nearly eight hours.
GitHub officials have said they are migrating from their own datacenters to Azure to try to ease some of the capacity issues. Daigle said it’s not a simple lift-and-shift process; rather, it’s a major re-architecture for the platform.
GitHub is roughly 60% through its Azure migration as of early October, he said.
It’s working with Azure storage teams on separating compute and storage for the Git distributed version control system to help alleviate bottlenecks caused by developers and agents working concurrently in the same repositories. Developers and agents made 7.38 billion commits on GitHub in September alone, according to the company.
GitHub’s recent growth would not have been manageable without Microsoft scaling expertise and Azure support, Daigle said. GitHub is in a unique position of being able to request major additional capacity, such as millions of CPUs, and work with Microsoft teams to provision it, he said.
The GitHub app and agent store: GitHub got its start as a platform for source control and collaboration. But its role is evolving as part of its Microsoft integration. In addition to developing and supporting Microsoft’s most successful Copilot, GitHub Copilot, GitHub has an important place in Microsoft’s agent-centric strategy.
GitHub is becoming “the store for anything that needs to be coded and needs some verification,” Daigle said. “We’ll connect you to whatever the right tool inside of Microsoft is for you to run that or whatever tool you’re using, not just (from) the Microsoft product suite.”
In the new world order, GitHub isn’t just the store for apps. Increasingly, it’s also the store for agents, which means GitHub is acting as the developer layer for Microsoft’s “agent factory” strategy.
(Microsoft co-founder Bill Gates envisioned Microsoft as a “software factory” that could produce software at scale. These days, CEO Satya Nadella and team talk about Microsoft as an “agent factory,” helping customers create agents at scale.)
GitHub currently serves more than 200 million developers, plus an unknown but quickly growing number of agents.
Developers need to be thinking in new ways when it comes to agents, Daigle said. They need to consider ideas such as scaling application programming interfaces (APIs) for agents separately from humans, creating APIs designed specifically for agents, rewriting documentation and tooling so agents can consume GitHub efficiently, and treating agent access as a fundamentally different workload from human access.
In the longer term, Daigle said he expects AI to make software economically viable for smaller audiences, such as a family, an individual or a single team. The app-store model may change substantially in the coming years, as people create highly specialized applications. GitHub could help like-minded users discover software for specific interests, such as a narrowly focused household or community use case.
GitHub historically has considered anyone who has seen or touched code to be part of the developer community. Daigle’s vision is for GitHub and Microsoft to let both professional developers and newer builders create and run software and agents without needing to manage token costs, uptime, monitoring or operational overhead.
GitHub will have more to say about its plans at the GitHub Universe conference in San Francisco, Oct. 28-29. The event won’t focus on code generation alone; it will include information on tools covering the full lifecycle, proactive security approaches and how to determine whether software is working well after deployment.
“We’re looking to solve the core parts of the platform, like Git, like Actions, in a way that I hope will make people excited and feel like they can absolutely trust GitHub for the next decade to come,” Daigle said.
Someone bursts through the door and announces, “There’s pizza!” Nearby, a team has duct tape and cardboard holding its prototype together. Another is debugging a model that won’t detect their movements.
This is a common scene during hackathons, and for a lot of people they’re the place to learn to build.
During several recent hackathons, we followed dozens of participants from the badge line to final demos and asked them what keeps them coming back.
Two days of intense focus
Teams come up with ideas for robots, apps, websites, hardware, and more, then try to get prototypes working before time runs out.
These aren’t small things, these are big endeavors, and they’re going to try to do it in basically two days.
Kyle, hackathon participant
Yana describes it as two days of intense focus on one small experimental project that moves the needle forward. For Nithya, it’s a rare chance to work as hard as you can on something you’re proud of.
It is basically an invention marathon.
Mike, hackathon participant
You don’t need a computer science degree
For most of computing history, building software required years of specialized training. Not anymore: with AI-powered tools like GitHub Copilot, natural language is becoming a universal programming language, and anyone with an idea can start turning it into working code.
Lance’s team hit a problem at the very last minute. They needed a news API so people could fact-check the information they were getting.
GitHub Copilot was able to not only create a frontend in 30 minutes but also connects with the actual API itself and integrates it all together.
Lance, hackathon participant
When they’re paired up with someone that’s a CS major, or when they just go and use AI to start a prototype and convince others to join in on their idea, it’s giving this autonomy for people to do what they’re most passionate about.
Kyle
The pizza is free, so are the new skills
Call it a learning event, Jon says, and people might not want to go. But tell them there’s pizza, and they’ll show up anyway, and they’ll still end up learning.
Jon calls it the best educational experience you can get as a technologist. Sarvesh puts it more simply: it’s a better way to learn.
Eric calls hackathons a huge grind of continuous work and continuous debugging. Vaishnavi’s team spent part of the time fighting a model that couldn’t detect their movements. “Like, nothing is working,” Vaishnavi said at one point. Mayank’s team got there through trial and error, with duct tape and cardboard.
Hackathons are a safe third place. It’s okay to fail and experiment. You’re not going to get an F, you’re not going to get fired over it.
Mike
Friendships of practice
Niels once searched Google to find out who came up with the term “hackathon,” and his own name came up. He doesn’t remember coining it, but he does remember setting laptops down wherever there was space, sleeping on the floor in sleeping bags, and going straight back to hacking after waking up.
If you feel you have to do it and you don’t like the people around you, then you’re missing the point.
Niels, hackathon participant
Akankshya found everyone open to collaborating and never felt pushed to quit or leave. Shehmeer could turn to the person on either side and ask how to do something. If they didn’t know, they usually knew someone there who did.
My first hackathon, I met two guys that are now my best friends.
Dev, hackathon participant
Zach has a name for it:
When you meet someone who you bond with over a shared interest, that’s a friendship of practice.
Zach, hackathon participant
Impact that changes lives
Mike’s first hackathon changed the course of his career:
At my first hackathon, I learned more in one weekend than I had in my entire university degree up to that point.
Mike
The next day, Mike looked in the mirror and told himself to forget being a lawyer and become a hacker.
Lee has seen hackathons act as a bridge between the spark of an idea and something rolled out into production. Shehmeer says organizers are trying to show a new generation that building isn’t out of reach.
For Yana, seeing the work happen pulls more people in, because there’s now a bit more of an open door.
You can solve problems yourself, you don’t have to ask someone else for permission.
Zach
Get started today
Everyone is welcome. Student or professional. All identities. All abilities. All backgrounds. All experience levels.
AI systems are gaining more autonomy while governments, companies, and researchers are still working out how much oversight they need. On the latest episode of This Week in AI, we covered the US debate over AI governance, new frontier models, persistent agents, world models, and practical uses for AI in healthcare and disaster response.
AI oversight is moving beyond company promises
The Trump administration announced a voluntary agreement with major AI companies that calls for internal safety monitoring, external audits, and independent board reviews. Because the agreement carries no legal enforcement, it raises a familiar question about how far self-regulation can go when companies are developing increasingly powerful systems.
Government agencies are also testing what existing law can do. The Federal Trade Commission launched an investigation into OpenAI, Anthropic, and other AI companies focused on potential consumer risks. Cases like these could help establish whether current consumer protection laws are enough or whether AI will require a more specialized regulatory framework.
For technical leaders, regulation can influence how organizations evaluate models, document risks, manage access, and choose vendors. As AI moves deeper into business processes, teams will need governance practices that can withstand outside review.
AI systems are taking on longer, more autonomous work
OpenAI just released Dots, its “proactive assistant” designed to retain context, work across applications, pursue multiple goals, and act independently. To do that work without waiting for a prompt, Dots needs standing access to the apps and data it works across. That also increases the amount of personal data an agent can reach and raises the cost of mistakes or misuse.
Google is also extending how long a model can work on a task. The company says Gemini 4 Argon can generate up to a million output tokens in a single response, far beyond the typical output limits of current frontier models. The goal is to let a model stay with long multistep work such as extensive coding or financial and legal analysis. Longer-running models and persistent agents aren’t the same thing, but both let AI finish more work without handing control back to a person.
As agents act more on their own, they need to anticipate the consequences of their actions. That becomes even more important when AI moves beyond software and begins acting in the physical world. World models aim to teach AI how these environments work, including cause and effect, which is why many researchers see them as building blocks for robotics and physical AI. World Labs (recently acquired by AMD) is developing spatial intelligence models for interactive 3D environments, while British startup Worldmodeldata has licensed nearly 1 million hours of video game data paired with player actions. Researchers from NVIDIA, MIT, and Oxford also introduced Physis-Lang, which uses descriptions of physical causes and effects to help video models learn why events happen. Researchers still don’t know which training approach will work best, so they’re testing several kinds of data and model design.
Together, these examples show how AI can be used for good by helping people work through complex information faster and, ultimately, save lives, while keeping qualified experts responsible for the final decisions.
What’s next
As AI takes on more work and enters higher-stakes settings, organizations will need clearer answers about access and accountability. They’ll also need to decide where human judgment remains necessary as systems become more capable.
Join us again next Monday for another episode of This Week in AI, when we’ll dive into more of the news, issues, and key developments shaping the AI era. And check back each Friday for the latest episode, or watch on YouTube, Spotify, Apple, or wherever you get your podcasts.
In my teamin Infobip, we handle message sending and serve as an endpoint for OTT providers such as WhatsApp, Viber, and Apple. Because of that, message sending is our core responsibility, and we need to be aware of any bottlenecks.
Initially, our service was measured at around 1,000 requests per second, but we didn’t know why.
From time to time, this raised concerns in the team: we should understand better where the bottleneck is and what causes it. Once we know that, we can make informed decisions to address client needs – for example, answering whether we can handle the traffic they expect to send. That was partly what led to this investigation.
How to run the tests
Application
Initially, with no experience in performance testing, I viewed the service as a set of methods – essentially a stack of calls. I assumed I could analyze it with a profiler. I also knew I shouldn’t run it on my computer, because other applications would affect the service’s performance.
That turned out to be a bad assumption. I’d suggest starting by measuring the application as a whole to get a broader view of its performance.
Integration
Since our application forwards messages further down the infrastructure and ultimately to the provider, it had to be adapted so that it would not send real traffic, while preserving as much of the original logic as possible.
Initial tests of this version showed that the service could handle a request in 0 ms, which seemed unrealistic. I realized that I needed to introduce a delay based on production, where the rounded average request processing time is around 50 ms.
Initial measurements
Having established that the service handles one request in 50 ms, we can estimate that it can process about 20 requests per second. Since the service is Jetty-based, it has roughly 200 threads available by default, which gives a theoretical peak of around 4,000 requests per second. With that in place, the service can now be tested.
Testing environment
DO NOT run such tests on a local machine.
Dedicated machine on our testing env was created for it.
K6 was chosen as test suite.
Rundeck to automate and parameterize remotely run tests.
Endpoints to be tested against:
/mapi-mock/1/dummy – performs Thread.sleep(50) – so it is blocking code run on Jetty thread pool
/mapi-mock/1/dummyReactor – also performs Thread.sleep(50)but publishing result on boundedElastic thread pool – it is still blocking code, but in that case one thread of reactor’s pool will be used
/mapi-mock/1/messages – non-blocking code which sends request through WebClient to real service
/mapi-mock/1/messagesReactor – non-blocking code which sends request through WebClient but publishing it on boundedElastic
JVM tuning
Make sure the correct memory parameter is used. In my case, HEAP_PERCENT was set to 60, meaning the JVM heap used 60% of the machine’s RAM. That allowed me to focus on machine RAM only, and the heap size would be adjusted automatically. This is not ideal, but it was good enough for our tests.
Metrics gathering
Initially, I ran the tests for 40 seconds, but I forgot that the metrics were collected every 30 seconds. That explains why I didn’t see any issues in the service metrics during the first run. When I extended the test duration to 180 seconds (a value I borrowed from similar tests) everything became visible.
You can also observe the machine directly with tools like htop or top, or any other tool that shows current CPU and memory usage. In the end, you can use whichever tool helps you complete the task.
One obvious but important rule: do not change two variables at the same time unless you know what you are doing. If you do, you won’t know which parameter had the effect. Change one thing, test it, then change the next thing, and repeat until you understand what is happening.know what is going on.
Gathering measurements
Initially, I decided to send up to 5k requests toward the service – it is below theoretical machine efficiency, but higher than what was experienced in the past. Lower CPU and RAM measurements are somewhat weird: 4 GB of RAM is not recommended as a base machine setup, because there is less than 2 GB of RAM left for the operating system, where not only our service lives, but also supporting applications. But both measurements for 2 CPU showed that it is way too low for handling any traffic efficiently and it showed quite high CPU and RAM usage, meaning that it could be improved. I finished the tests with 5k requests/s, when performance was quite near predicted performance.
The theoretical peak was reached, so it was time to check whether it could be improved. I tried switching from Jetty pool to Reactor pool, but it turned out (my bad for not remembering Project Reactor documentation by heart) that to reach the initial performance I had to configure the machine with 20 CPU, because Reactor constructs the pool to have 10 * CPU threads. So when I matched theoretical performance again, but now with the Reactor pool, I thought about what could be done next to improve. The last thing to check was how non-blocking code would perform, so I started using the /mapi-mock/1/messages* endpoints for testing. As one would predict, it was hit. The application peaked at 7,000 requests/s, but with quite large demands for hardware. So the last thing to check was to try to find reasonable settings to have big enough performance with not much resources assigned.
Now we are focusing on the non-blocking endpoint, so it more reflects the real-case scenario, and on limiting resources. The first thing that is apparent is that when it comes to non-blocking code, there is no difference between Jetty thread pool and boundedElastic thread pool (the latter results in some queuing because the rule of 10 * CPU threads is always true). We can see performance downgrade with CPUs removal, but not apparently linear. I had to say “stop” somewhere, and I decided to stop with 8 CPU and 8 GB of RAM, for which we received ~3.5k requests/s. Good enough with not much resources assigned. As we can scale horizontally, adding more machines will use more resources, but give more resiliency back.
The thing with VT…
As you can see, I focused on measuring reactive vs non-reactive code.
It was our go-to architecture at that time, while VT was being adopted slowly (we were afraid of thread pinning, so we waited until it was addressed in JDK 24 to start adoption). After I published it internally, I was asked specifically about VT – how the app would perform when configured to use this not-so-new-but-fixed feature.
It turned out that VT gave the app a boost. With optimal configuration (8 CPU and 8 GB of RAM) the app was able to handle 10k requests/s. Why is that? Any code, regardless of whether it is blocking or not, will be run on a virtual thread, which ideally should not block a system thread, leading to the case where system threads are used only for real logic and not for waiting. VT can improve applications where there is a lot of blocking code. In the rest of the cases, it is “just” syntactic freedom, no more looking at which operator do I need?
The outcome
When I performed the tests, I observed that just switching to boundedElastic (i.e. “use Reactor and you will get higher TPS”) didn’t increase throughput – it actually reduced it. Not remembering the Project Reactor documentation by heart, I was surprised, but after some research I found that boundedElastic provides 10 * CPU threads, which gave us a total of 60 threads to service all requests previously handled by 200 Jetty threads.
When the CPU count was adjusted to match 200 Jetty threads (20 CPUs) we reached the expected throughput. But it was wasteful in terms of resources. In the Reactor world, people often say that blocking is bad, and my service was actually doing exactly that, on purpose. I thought it was time to switch to a different endpoint.
So I was able to nearly reach the estimated performance of the service, which was good. That allowed me to move on and investigate other ways of improving it, such as using non-blocking code or VT.
What I’ve learned
It is always about finding the bottleneck and finding out if we can improve it, for example by adding resources or redesigning it. If it is not to be improved, then this is the peak capacity of the service.
Sometimes questions from clients can actually help. It is just like with a new person joining the team: they will ask obvious questions about all-known solutions, but if you want to respond accurately, you have to dig deeper, and sometimes you find out that this simple question leads to an investigation and even to questioning whether it is correct at all.
Also, be aware of architecture. In that scenario, we tested a service which accepts requests, so it is one of the first applications, or APIs, which the client experiences. Being first means that there is the rest of the system, which can in turn have throttling applied. What does it mean? Even if we accept 10k requests/s, sometimes we cannot send them immediately, because an independent system cannot handle, for example, more than 2k requests/s. So we can be fast for clients, in the sense that we accept their traffic as fast as they want, but we should always remember that this is not the whole picture, and the client should be informed about it.
In my test there was no DB, just API-to-API calls. Having any limited, and even worse blocking, resource like a DB in the way of processing a request will definitely impact its performance.