Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
161021 stories
·
33 followers

Bored at Work but Exhausted? That’s Still Burnout

1 Share

Nothing was wrong when Chris came to me.

Good pay. Good life. Ten years at the same company, senior engineering leader in space and deep sea robotics — no crisis, nothing broken. He’d mentally mapped out four more years on autopilot and figured he’d sort out what came next once he retired.

Here’s how he put it: “I had good pay, a good life, and I was enjoying it — but I wasn’t getting much out of it.”

That’s still burnout. It just didn’t come from doing too much.

Two and a half months later he’d left the job he’d held for ten years, moved across the country to be closer to family, and started as an engineering director. Nothing had to break first. He simply got clear on what he actually wanted, and a four-year plan collapsed into now.

I’ve heard hundreds of engineering leaders describe some version of what Chris felt, and almost all of them apologize for calling it burnout. The pay is good. The reviews are good. Nobody’s pulling 70-hour weeks anymore. What am I even complaining about?

Here’s what’s actually happening.

Every role demands a certain level of skill from you. When you’re new, your capability climbs fast to meet it — steep, uncomfortable, exciting. Then you get good at the job you were hired for, the new inputs stop, and the curve flattens.

Stay on that flat part long enough and it doesn’t hold you level. It pulls you back down. Another project like the last project. Another hire like the last hire. You feel like you’re losing your edge, because you are.

The most exhausting run I’ve ever taken down a mountain was the easiest one — a flat cat track in Park City that went on forever and left my calves burning. Ten inches of powder the next day, black diamonds all afternoon, and I had energy for hours. Easy terrain drains you. Your edge gives it back.

You’re not tired because the work is too hard. You’re tired because it stopped asking anything of you.

One caution before you go any further with this: depleted and stuck feel almost identical from the inside, and they need opposite responses. If you’re running on empty from an unsustainable workload or a culture that’s grinding you down, what you need first is recovery, not a new challenge. This episode is for the other version — the one where you have plenty left in the tank and nothing worth spending it on.

If that’s you, the real question is which one you’ve outgrown: the role, the company, or the direction. In this episode I walk through how to tell the difference, the one question that decides whether you push inside your company or pivot out of it, and three small experiments you can run instead of waiting for clarity to arrive on its own. The free Next Move Experiment below has the full diagnostic plus all three, so you can work it through on paper before you make any giant career decision.

You are not stuck with the career you designed five years ago.

The Happy Engineer Podcast

Bored at Work but Exhausted? That’s Still Burnout

LISTEN NOW

Back to ALL EPISODES

 

What You’ll Learn

 

Depleted and Stuck Feel Identical — They Aren’t

They produce the same flat, tired feeling, but they come from opposite causes and need opposite responses. Depletion means the tank is empty — unsustainable hours, a draining culture, no recovery time — and the fix is rest, boundaries, or support, not a new challenge. Stagnation means the tank is full and nothing is asking you to use it. Confusing the two means either running an experiment you don’t have the energy for, or resting when what you actually need is a bigger problem to solve.

 

The Growth Curve Explains Why “Nothing Is Wrong” Still Feels Like Burnout

Every role demands a certain level of skill. Early on, your capability climbs fast to meet it — that’s the steep, exciting part. Then you master the job, the new inputs stop, and the curve flattens. Stay on the flat part long enough and it starts pulling you backward instead of holding you level. Chris lived on that flat part for years with nothing visibly wrong, and it was still costing him energy every day.

 

The Push-or-Pivot Question

Once you know it’s stagnation, the next question isn’t whether you could survive another year where you are — it’s whether the company you’re in can actually take you where you want to go next. Impact, scope, learning, autonomy, money: name what you want more of, then ask honestly whether this vehicle gets you there. The answer decides whether you push for more inside your current role or start looking outside it.

 

Three Experiments, Not One Big Decision

Clarity doesn’t arrive from thinking harder about the question — it comes from evidence, and evidence comes from action. Instead of trying to solve “role, company, or direction” all at once, pick one small experiment tied to whichever answer feels most true and run it this week. Chris didn’t get handed a new career: he got clear, and a four-year plan collapsed into now.

 

About Zach White

Zach White is the founder of Oasis of Courage and the host of The Happy Engineer podcast, now in its second season with more than 200 episodes published.

He holds mechanical engineering degrees from Purdue and the University of Michigan and spent his corporate career as an engineering leader at Whirlpool before leaving in 2019 to coach full time. Since then he has coached more than 300 engineering leaders at over 200 companies — big tech to rust belt, and everything in between.

His flagship program, the Lifestyle Engineering Blueprint™, is a 12-week coaching program that helps engineering managers and directors get promoted without handing over their nights and weekends to do it. He also runs an invite-only Mastermind for Blueprint graduates and hosts OACO Live Events around the country.

He built all of it around one belief: you shouldn’t have to sacrifice your life to reach your potential at work.

 

Links & Resources

Free download — The Next Move Experiment
A diagnostic that separates depletion from stagnation, the push-or-pivot question, and three experiments to run this week — pick one and get real evidence before you make a giant career decision.

Watch next

Work with ZachBook a free Career Growth Audit 

ConnectFollow Zach White on LinkedIn

 

FULL EPISODE TRANSCRIPT:

Please note the full transcript is 90-95% accuracy. Reference the podcast audio to confirm exact quotations.

[00:00:00] Burnout, it doesn’t always mean that you’re doing too much or that you’re sleeping too little. In fact, sometimes you’re just tired because you’ve stayed too long in a role doing the same hard things for so long that they’re not really hard for you anymore.

Expand to Read Full Transcript

I’ve hit that point more than once in my own career, even as a coach, and nothing was wrong on paper. But you feel frustrated. I felt flat and wondering then if something’s wrong with me. And there’s a reason that happens. So today I’m gonna show you what that really looks like.

And there’s, there’s 3 experiments that you can run to get your career back on that growth edge. But before we go to the whiteboard and model this, before I share these experiments with you, we need to go back to the fall of 2002 when I was playing junior varsity tennis. For my high school. Now, every athlete knows that you rise to the level of your training and your competition. And if you want to improve as an athlete, you need to level up your coaching, level up your teammates, level up your training partners, and play against better opponents.

If you keep training at the same level, playing against the same people, you stop improving and In fact, you might even catch yourself losing games to opponents that in the past you’ve beaten easily. Well, in 2002, I was playing junior varsity tennis for my high school. And at the beginning of every practice, we would pair up and volley with someone on the team to warm up and get ready for practice. And you’d generally pair with somebody who was at your same level.

You know, the best players would volley with each other and the worst players would volley with each other. Well, I was kind of a middle-of-the-pack And if you watched me and my buddy Joel warm up, it wasn’t exactly Wimbledon, okay? I’d put the ball in the net 2 or 3 times in a row, send one soaring past the baseline, and just kind of write it off as, hey, we’re warming up. But one day I got to practice early, and the only other player who had arrived was Steven. Steven was our number 1 varsity singles player. He was exceptional. He was on track to play college tennis on scholarship. And Steven looked at me and said, hey, let’s warm up. And I thought, you gotta be kidding. You know, I’m not nearly good enough to volley with Steven.

And I was kind of freaking out. But what happened actually shocked Steven just as much as it shocked me. As we started to volley, you know, Steven puts a lot of pace on the ball. Well, I struck the ball better than I maybe ever had in my entire life and started to build some confidence. I started to get really excited. It felt incredibly fun. I was putting more pace on the ball. I was putting more topspin on the ball. It was some of the best tennis I’ve ever played in my life. And it lasted about 10 minutes. Then the number 2 singles player showed up and he kicked me off the court so he and Steven could warm up together as they always did. And I went back to playing with the other middle-of-the-pack players. And the truth is, it wasn’t quite as fun. I wasn’t anywhere near as engaged with playing the game because I had tasted what it was like to be at my edge, to play at my best. And I couldn’t draw that part of myself out when I was volleying with people who were Constantly putting it in the net, and the same thing can happen in your career. I call this phenomenon burnout from boredom, and I’ve heard it described by hundreds of leaders who I’ve talked to.

And I want to show you what actually happens. You think about in your career what you’re developing are capabilities, traits. We’re going to bundle it all underneath skills. What you can do, your capabilities, the results that you’re able to drive, you as a leader, the impact you can make. You’re developing skills, and those skills develop over time as you do new roles. Well, the role that you’re in has a level of skill that’s required for you to perform well. So let’s just put a line on this chart where the y-axis is the development of your skills and The x-axis is time, developing skills over time. And you can probably relate to the curve that happens when you’re new in a role, new at a company. It’s really exciting. And you have this steep section of growth, a steep section of growth at the beginning. This is the fun part. It can be challenging. You’re drinking from the fire hose. Boom. But it’s that kind of challenge that is what you signed up for. And when you’re growing and you’re being developed and you’re getting all of these new inputs, that’s a really exciting time in the role. But then you know what happens as you approach the skills that are required to do your job.

You start to flatten out a little bit, right? You get this flatter part of the curve, and the reason that the flattening happens. Is because once you can do the job that you’ve been hired to do, you stop getting all of those inputs. You stop getting all of those development opportunities because the role is designed for you to perform at a certain level. And so you start to flatten out, and that’s okay. That’s normal. In fact, for some of us, it’s a bit of a relief at first because the pressure of that steep growth curve is taxing. It takes a lot of energy, but Here’s the part that we don’t like. I’m zooming in on this. As you pass the required development line, you know, especially as a top performer, you want to stretch this even further past, right? You want to grow beyond what’s required for the role. But there’s this tension that starts to build between what you can do and what the role demands from you. And that tension It actually has an effect of pulling you back down a little bit. You start to notice yourself not just plateauing, but almost feeling like you’re regressing. I hated this feeling, the feeling that not only had I stopped growing in a role, but that I was losing my edge, that the boredom of just doing another one of those.

This project, yeah, it’s a new project, but it’s another project just like my last one. This team member, yeah, they’re new, I just hired them, but it’s onboarding another team member just like any other team member. And that growth curve not only flattens, but it’s like the demands of your job, it pulls you down to that level and it won’t let you escape to something more. Have you ever felt this before? That feeling that I’m just trapped at this level? And here’s the worst part, it’s usually a pretty good Situation. It’s an okay life. It’s a career where you’re making good money. You might even be getting excellent performance reviews, but you’re not at that growth edge. And there’s something about being stuck in that plateau, in that feeling of stagnation, that can be even worse than being let go. At least then you know exactly what to do. Go find a new opportunity. And guess what? You get to go right back. to this exciting growth curve section of learning something new at a new organization. It lights you up again. But when you’re trapped in that flat spot, oof, that is a place that can leave you feeling really empty and drained. It’s burnout from boredom.

I actually had a client recently go through my Lifestyle Engineering Blueprint coaching program. His name was Chris Yonker, and Chris left a video on our website explaining his own experience with this type of burnout. He’d been 10 years at the same company, good pay, stable life. He was not in any type of crisis, but mentally he sat there and thought, well, I’ve got, you know, 4, 5, 6 years of career that I want to continue, and how can I make the most of this? But that feeling was muted by this plateau. And I wanna give you the direct quote, exactly what Chris said, because it’s exactly what you might be experiencing if you’re in this place. Chris was a senior engineering leader in space and deep sea robotics. Here’s what he said. 2 and a half months ago, I came to Zach because I was feeling stuck in my career. I didn’t know where to go. I had good pay, a good life, and I was enjoying it, but I wasn’t getting much out of it. You see, that’s that growth curve flattening. It’s like good pay, good life, but it just, it just wasn’t feeding my soul.

I didn’t want to go in and do it. Back to Chris. He said, hey, look, I’ve got about 4 years left and I could retire from my current company. So how do I make the best of those 4 years and then just figure out what I want to do going into retirement? Oh, it’s that feeling of like, can I just get through this next season? And maybe you’re not near retirement like Chris was, but You still think about, man, I’ve got all this work ahead of me. Do I really want to just be a manager for the rest of my career? Do I really want to just do more projects as a senior engineer for the rest of my career? Well, back to Chris. He said, I got really clear on my vision, really clear on my purpose, understanding what I wanted in life. Do you understand what you want in your life? Accelerating those 4 years to today, I’ve been able to make some really big breakthroughs, and I actually just left my job of 10 years, moving across the country to be closer to family and start something brand new as an engineering director next week. Back to the growth curve.

You see, he needed to get back to that place, that place of challenge, that place of enthusiasm and excitement. When you stay out here in this flat spot too long, It just doesn’t matter how good it might be on paper. We have inside of us a compelling desire to be at that growth edge, and that’s what I want for you as well. For Chris, he thought he just needed to ride it out for a few more years to retirement. But what he realized when he got clear in terms of his vision, his values, his purpose, what he really wanted as we worked through the Blueprint journey together was that no, I want to continue to operate in a way that lights me up where I can perform at my best. And that means getting to our growth edge. Let’s talk about how you can start running experiments to activate that in your own career.

And by the way, if your career does look good on paper like Chris’s did, this can be a hard problem to even admit. You make good money, you’ve got the stability, maybe you’re Not working long hours. You’ve got 40-hour work weeks. Like, all of the things are good. And so you get that sense from the inside. It’s like, what am I even complaining about? But feeling stuck is still painful, and it’s information you need to respond to. So I made the Next Move experiment to help you figure out whether you’ve outgrown the role, the company, or maybe the direction of your career entirely, and then run one small experiment before you make a giant career decision and change course in something that you may regret.

So I’ve put all of what we’re about to walk through into a simple doc that you can download and work through. Grab it in the link in the description and follow along as we walk through the next move experiment. The question I get when someone is in burnout from boredom is, do I need a different role? Do I need a different company? Or in an extreme case, do I need to leave engineering altogether? Maybe start my own company, do something completely in a different direction. And here’s the trap. You can spend months trying to think your way into the right answer for that. But when we’re up here in our mind, Clarity is not going to arrive. Clarity does not arrive before action. Action is what will lead you towards clarity.

So let’s talk about the Next Move experiment. There are 3 different kinds of experiments that we can run. The first is a role experiment. And I’m calling these experiments because I’m not telling you that you need to change your role. What I’m telling you is that you need to take an action now to help you understand with clarity if changing roles will be the right next thing for you. We can run a role experiment. Should I move from this type of work to that type of work, from this team to that team, this department to that department? There are so many ways that you can activate growth again by changing role within your existing organization. But the second kind of experiment is the bigger one lots of people want to run. Is it time for a new company? A company experiment. Hey, Zach, I’ve had this conversation before. I’ve talked to my boss about my desire for growth, my hunger for the bigger project, my interest in learning new things, and nothing is happening. I’m just not getting the support here. Or, Zach, I’m going to have to wait for my boss to retire before there’ll ever be an opportunity here. Okay, maybe a company experiment is the right place for you to take action next. And in some cases, you may have reached a point in your life and in your career where it’s time for an entirely new direction. Just had an amazing conversation With a past client named Hadi, who’s had some tremendous success.

And in fact, the vision that we wrote together in our Blueprint coaching has come to pass. He has actually overdelivered and crushed it on the things that he set out to accomplish. And now he’s thinking about what’s next. Is it time for a new direction? You might be thinking those same things, but we don’t want to disrupt an amazing career that we’ve built with a move that we haven’t tested. Let’s Run an experiment. 3 kinds of experiments you can run: a role experiment, a company experiment, and a direction experiment. So how do you know which type of experiment to run first? Good question. The first thing I’d ask you to consider is, do you feel depleted or do you feel stagnant? And I know those 2 words might feel similar, but let me explain the distinction. Sometimes you’re energetically depleted. You have been working 60, 70, 80 hours a week. You’ve been in a toxic, terrible culture.

You’ve been micromanaged. You’ve been treated in a way that is emotionally taxing and your energy is depleted. There’s nothing in the tank, and you’re experiencing the kind of burnout that feels like an empty rock bottom depletion. That is different from the feeling of simply no longer growing, the burnout from boredom that we’re talking about here. And these 2 require different responses. If you’re in a depletion state, these experiments It may not work for you at all. Depletion requires a different intervention. But if you’re in stagnation, then we can look at these 3 and say, which way do I want to go? How can I find that growth edge again? The next important question: is it time for me to push or is it time for me to pivot? This is the bigger question when looking at stagnation, because our temptation is to just run away from the situation. I don’t know how to solve it here. I’m in this plateau, so I’m just going to leave. I’m going to leave the team, leave the company, leave the direction, whatever that is. Well, I want you to ask yourself, can I imagine achieving the career and the life that I actually want? From inside the company that I’m in today.

Here’s a more general way to think about it. When we do coaching together, we build a really clear vision and we create a clear set of values that drive you towards that future. So if, if your career is the vehicle, it’s the, you know, it’s the Ferrari that you’re going through life in, and your vision is in the driver’s seat and your values are in the passenger seat. Mm-hmm. And you’re asking yourself, can I get to the destination in this vehicle? Can I achieve the career and the life that I want inside the vehicle that I’m in today? And if you can see yourself at full speed, if this vehicle were to continue to the peak that it could, to its max capacity to deliver in your career and life, could you get to your vision? And if the answer is, Yeah, I think it could. I think that there are opportunities inside my organization that look like where I want to go. I do think that my values can be aligned and honored and work for me in this company. Well, then great. Let’s continue to push because changing vehicles comes at a cost. Changing vehicles means you’ve got to ramp back up and build speed again. So if you believe that the vehicle you’re in could take you where you want to go, then let’s focus in on a role experiment. Let’s use the existing momentum we have to look carefully at what can I do to push to a growth edge with a role experiment? If the answer is no, Zach, this company, this organization, this direction, I don’t see my vision ever coming to pass here. I cannot even imagine.

A situation where my vision could happen here. If that’s true for you, not because you’re feeling a little unhappy, not because you’re upset at a boss, but because you’ve taken a really clear, open, and honest look. Zach, it’s not gonna happen. Okay. It’s time to pivot, and we need to make that decision. Is it a company or a direction experiment? Now, how do we decide which one of those 2 is right for you? Is it a company or direction experiment? Here’s the simplest way to think about it. Fulfillment in our work, fit between your job, your career, and yourself comes down to what you are doing all day, every day. The actual activities that fill your calendar, the work that you show up to do. And making money is great, but if you show up to work and what you’re doing from 8 to 5 is actions you just don’t enjoy taking, well then we’ve got a real problem in terms of direction. Right? That’s not going to change if you just switch companies, right? You’re going to be doing similar work just in a new environment with a new opportunity, with new scope, new trajectory.

So the company experiment is great for you to run if you enjoy this kind of work. You do see yourself continuing on the path. It’s just that the place right now looks like a vehicle that won’t take you to your vision. If that’s you, a company experiment is great to run. But if you can say honestly, I just can’t see myself Waking up and doing this every day for the rest of my life. My vision is something completely different. It’s time to go run the direction experiment. And I need to be honest with you, most people think they know what they want, but at best they have a vector and not a vision. They know they want to lead people or make a bigger impact or make $1 million a year or whatever. And it’s these general things that you know that you want. But general things, a vector towards the future, is not the same as a clear vision of the future. And if you don’t know the difference and you don’t have a vision, then you need to reach out. We need to talk about that because none of these experiments are going to be clear and helpful if you don’t have that piece first.

But let’s get back to where we go from here. You’ve got an experiment that you want to run. Why is this so important and how is it gonna help you get back to the growth edge? Remember, we’re engineers, so experiments are designed to test hypotheses, to answer a particular question. And so if you’re running a role experiment, the question that you’re testing is, can I get back into meaningful growth without leaving my company? That is a role experiment. I’ll give you 2 or 3 really great options for exactly what to do in the free download in the description. I won’t go through all of those now because you’ll want to do some critical thinking about which one is right for you. When you move on to a company experiment, the question here that you’re testing is, is my career actually wrong or Or is the environment wrong for the career that I want? And if that isn’t where you head, if it’s a direction experiment, then the question is, have I outgrown or evolved into something that doesn’t match this path at all? Have I outgrown the path itself? Those questions, we wanna take specific actions to test.

And I’m gonna, like I said, provide for you 2, 3, 4 different ways that you can run those tests as an inspiration. And I want you to go through this tool and I want you to ask yourself, following the guideline, which test, which experiment is for me? Which specific action is the right one for me to take next? And most importantly, how quickly can you get into action on that experiment? Last piece. Experiments in career are best run with support outside yourself. So don’t go alone. Get a coach, get a mentor, get a colleague who you trust that you can work with. I think about taking a big group of my clients out to Copper Mountain, Colorado for a live lifestyle engineering event. We did some amazing coaching. We had some amazing fun. on the slopes together, and I love to snowboard. And as a snowboarder, it’s just like we’re talking about here in our career. You begin on the bunny hill, you begin with no skill at all, and you’re constantly falling down, and, you know, your knees and your butt are pretty sore at the end of the day. But once you master those beginner slopes, you want to move on, go from greens to blues, and eventually blues to blacks, and blacks to the big Double black bowls on the back of the mountain. And the point is, if you don’t continue to move up and find your edge, yeah, it’s always fun, but it’s not as fun.

And I remember being in Park City, Utah on a cat track that lasted forever. It was like the longest, most boring run of my life. And at the end of that, I was genuinely tired and fatigued. My, like, my calves were burning and my body was just sort of stiff and sore from this really long, boring run. And it took a lot of energy out of me to do this incredibly simple and boring run. And I think it’s as much mental as it is physical, because the next day we got 10 inches of powder. And my buddy Jack and I went out to the back bowls and we just did run after run.

And I had endless energy. And it was incredible because in those more exciting and difficult conditions, I was able to find my edge again. It’s time for you in your career to get yourself back to that edge, your growth edge, where you can get out of stagnation, live into your full potential, and disrupt the feeling of burnout from boredom. So like I said, I made the Next Move experiment to help you figure out whether you’ve outgrown the role, the company, or maybe the direction that you’re headed entirely. And then run one small experiment now before you make some giant career decision, because you will not find clarity by sitting and thinking. Grab it from the link below. And if you want to go further with these ideas, I’d really encourage you to check out a video that I made recently that is going to explain another one of the main reasons that stagnation can happen in your career.

If you’re feeling stuck, if you’re not sure, it’s called Why Great Engineers Get Passed Over for Promotion. We’re gonna put the link to that video here. If you’re feeling like, yeah, Zach, I am growing, I am doing great work, I don’t think it’s this thing you’re talking about with burnout from boredom, That video may have the exact key to exactly why you’re getting passed over and someone else who’s not even as good as you is getting that next opportunity. So go check that video out next. I’ll see you there. Let’s do this.

 

Back to ALL EPISODES

The post Bored at Work but Exhausted? That’s Still Burnout appeared first on OACO.

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

How and Why We Built an AI Assistant

1 Share
Directions on Microsoft built its own AI Assistant to help with Microsoft technology and licensing queries. Analysts Barry Briggs and Andrew Snodgrass share experiences and lessons learned with Mary Jo Foley.



Download audio: https://www.directionsonmicrosoft.com/wp-content/uploads/2026/09/season5ep14atlasai.mp3
Read the whole story
alvinashcraft
16 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it.

1 Share

This week, Mistral announced it raised €3 billion in a Series D funding round, pushing its post-money valuation past €21 billion. With the influx of cash — $3.5 billion in US dollars — it plans to expand its frontier research, scale compute capacity for model training, and grow its infrastructure. 

Where Mistral’s allocating new funds suggests what the French AI company is betting on for the future of AI power: open-weight models can only do so much if the compute and infrastructure underneath remain concentrated among a few key players. 

Open weights can only go so far

So far, model superiority has been a major factor in who gets to rule the AI roost. Some open-weight advocates have been touting open-weight models as a way to combat this concentration by giving developers more choice over the models they use — and a way to escape dependence on proprietary APIs. This way, rather than relying exclusively on one provider’s model, developers can adapt open-weight models for their own use.

The catch? Running powerful models takes enormous amounts of compute. Training frontier models — and serving them at high volume — requires compute capacity concentrated among a relatively small number of labs, chip suppliers, and infrastructure providers.

For his part, Dario Amodei, CEO and co-founder of Anthropic, challenged that vision for open-weight models last month in an exchange on X, where he wrote that open weights “are nowhere near a sufficient solution because they simply shift the concentration somewhat to those with the most compute and chips.” 

But the expansion plans Mistral briefly outlines in its funding news suggest there’s a different way to combat that dominance: Don’t stop at opening the model. Build more of the stack, instead. 

So Mistral is building more of the stack

“Mistral is the only AI company in the world building the full stack required to answer that question,” claims the French AI company, writing about how organizations can take advantage of AI for mission-critical needs without giving up control of the infrastructure and intelligence loop.

For Mistral, building that stack means developing open-weight models and the infrastructure and compute capacity on which those models run, along with the downstream products that bring them into production. And with a new €3B in the bank — led by Samsung Electronics, with Scaleup Europe Fund, managed by EQT, and existing investor PSG Equity as co-leads — parts of that stack will keep expanding. 

Looking ahead, Mistral says it aims to use its full stack and open approach to AI to free customers from dependence on a single vendor’s roadmap, pricing, and availability so they can build on its stack “without exposing their most valuable data, workflows and institutional knowledge to anyone outside their walls.” 

That addresses one piece of Amodei’s critique of open-weight models. Because Mistral’s stack includes not only the models but also the compute, infrastructure, and production layer, its open-weight strategy depends less on rival-controlled infrastructure. 

It’s been moving this way for a while 

Launched three years ago, Mistral has made a name for itself by releasing open-weight models. Interestingly, it’s also been expanding into the infrastructure layer as of late. 

Last month, the company said it would begin hosting third-party open models, putting the likes of GLM-5.2 from China’s Z.ai on the same infrastructure as its own models — another move that suggests it sees the infrastructure layer as an increasingly important part of the AI race. 

In July, Arthur Mensch, co-founder and CEO, Mistral, added to the case for more openness by taking to LinkedIn to express his concerns about dependence on closed-model providers, writing: 

“Of course you need to use open-source models if you’re an enterprise leader. Closed-model providers, that are now forcing data retention, are gaining immense leverage on your business if you don’t.” 

Bigger picture, it looks like Mistral’s betting that whoever ends up ruling the AI roost will need more than the best-performing model; they’ll also need to control enough of the surrounding infrastructure to give customers choices about which models they want to use and on what infrastructure. 

Whether this can meaningfully shift AI power, though, remains to be seen. 

The post Mistral wants open-weight AI to compete at the frontier. It just raised $3.5 billion to do it. appeared first on The New Stack.

Read the whole story
alvinashcraft
34 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Solution Colors: Telling Visual Studio Windows Apart at a Glance

1 Share

When several Visual Studio instances are open, identifying the right one becomes a surprisingly repetitive task. Two windows showing different checkouts of the same solution can look almost identical. Reading the title works, but a small visual cue is easier to spot.

That's the problem behind Solution Colors. Inspired by the Peacock extension for VS Code and a Visual Studio feature request, it associates a color with a solution or folder without replacing your editor theme.

Overlapping Visual Studio windows with purple, cyan, green, and orange bottom borders.

Give each solution a visual identity

Right-click the solution in Solution Explorer, choose Set Solution Color, and select a color. The predefined choices use the familiar document-tab color palette, and Custom... opens a color picker when you want something different. Open Folder workspaces are supported too.

Solution Explorer context menu with Set Solution Color expanded and Gold highlighted.

The color can appear around the window, behind the solution name in the title bar, and in taskbar thumbnails. Taskbar icon overlays are also available, with the documented limitation that taskbar items need to be ungrouped. You can choose the surfaces that help you and leave the others alone.

Under Tools > Options > Environment > Fonts and Colors > Solution Colors, you can adjust border placement and thickness or enable automatic color assignment. Automatic mode selects from the palette using a hash of the solution path; explicitly saved colors take precedence.

A small extension with two kinds of integration

The project is an in-process VSIX built with the Visual Studio SDK and the Community Toolkit. Its ToolkitPackage registers commands and listens for solution and folder open/close events. The menu lives in a VSCT command table, while the color commands share a generic BaseCommand<T> implementation.

That is the conventional part. Coloring the existing window chrome is less conventional.

The implementation searches the WPF visual tree for named shell elements such as BottomDockBorder and PART_SolutionNameTextBlock. It changes brushes and border thickness rather than introducing a new tool window. For the title label, it also uses reflection to access the foreground property and Visual Studio's ColorUtilities.CompareContrastWithBlackAndWhite to choose black or white text.

This is an important tradeoff for extension authors: finding an element in the shell's visual tree is not the same as having a dedicated, stable extensibility contract for it. Names and structure can change. The lookup code is worth isolating, and these integrations need testing against the Visual Studio versions you support.

Decoration must not get in the way

One revealing detail is hit testing. The extension disables hit testing on the non-top borders it colors so they don't intercept mouse input. But it deliberately leaves the top element alone: MainWindowTitleBar contains interactive UI, and disabling hit testing there would also block menu interaction.

That distinction is easy to overlook when the feature appears to be "just a border." Decorative changes still participate in layout and input routing.

Timing matters too. Package initialization does not guarantee that every shell element is ready. The startup path makes a bounded series of colorization attempts, with delays between them. UI changes switch to the main thread, while Git branch file reads are dispatched to a background task.

Taskbar integration uses a different mechanism: WPF's TaskbarItemInfo, noninteractive ThumbButtonInfo entries, and an overlay image. There is no need to reach into the taskbar's visual tree.

Branch colors are a state problem, not just a brush problem

The settings offer one color across branches, a separate color per branch, or a combined gradient. Assignments are stored as simple branch:color lines in a color.txt file, normally under the solution's .vs directory. An option puts the file in the root instead. The parser also accepts the older single-color format.

There is a compatibility detail worth knowing: the implementation uses the literal key master for its shared/base color. That is an internal convention, not automatic discovery of a repository's configured default branch.

A recent fix illustrates why separating decisions from rendering helps. In per-branch mode, a branch with an explicit color should remain colored even when no base color has been assigned. The code now expresses the removal decision in ShouldRemoveColorization, with focused unit tests covering the different modes and automatic-color behavior.

That is a useful pattern beyond this extension: put the state rules in a testable method, then let the UI code apply the result. You shouldn't need to launch an experimental Visual Studio instance just to verify whether a missing base setting overrides a branch-specific one.

Try it, or borrow the ideas

If you regularly juggle Visual Studio windows, install Solution Colors from the Marketplace and give your solutions a visual identity.

If you're building extensions, browse the source on GitHub. Start with SolutionColorsPackage for the lifecycle, ColorHelper for the shell integration, and the tests for the branch-color rules. A few pixels of color turn out to be a useful case study in extending the IDE without getting in the user's way.

Read the whole story
alvinashcraft
50 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Agent Gateway: The Next Evolution of the API Gateway

1 Share

For over two decades, the API Gateway has been one of the most important pieces of modern application architecture. Whether you’re building a mobile application, a SaaS platform, or a collection of microservices, chances are every request from your users passes through an API Gateway before it reaches your backend.

But something has changed.

The clients making requests are no longer just browsers and mobile apps. Increasingly, they’re AI agents, and AI agents behave very differently from traditional software. As a result, we are beginning to see the emergence of a new architectural component called the Agent Gateway.

To understand why it is needed, it helps to trace how gateways have evolved.

The API Gateway: A Reverse Proxy with API-Specific Capabilities

When I think about gateways in software, the first concept that comes to mind is a proxy.

A proxy is an intermediary between two systems. A forward proxy acts on behalf of a client, while a reverse proxy sits in front of one or more servers and acts on their behalf.

At its core, an API Gateway is a reverse proxy with capabilities designed specifically for managing APIs. That may sound like an oversimplification, but it is a useful mental model.

An API Gateway sits between API consumers and the services that fulfill their requests.

Imagine a user opening a food-delivery application and tapping Order food. The application sends a request through an API Gateway, which routes it to the appropriate backend services.

The gateway may handle routing and load balancing, authentication and authorization, rate limiting and quotas, TLS termination, Protocol translation, Request and response transformation, Caching, Logging, tracing, and observability, API versioning and lifecycle policies and many more.

Most importantly, the API Gateway makes a fundamental assumption about its client:

The client already knows which API it wants to call.

A mobile app doesn’t ask “How do I order food?”. Its developers have already encoded that logic into the application. The client knows that it needs to send a request such as:

The API Gateway’s job is to authenticate, govern, route, and reliably deliver that request at scale.

AI Agents Don’t Work That Way

Now imagine replacing that mobile app with an AI agent. Instead of receiving a predetermined instruction to call a specific endpoint, the agent receives a goal. “Book me a hotel near the conference venue under $250.”

To accomplish this, the agent might need to:

  • Identify the conference venue
  • Search multiple hotel providers
  • Compare prices and cancellation policies
  • Read reviews
  • Calculate travel time
  • Check room availability
  • Ask the user to clarify missing preferences
  • Reserve the selected room
  • Send a confirmation

The exact sequence is not necessarily known in advance. It may change depending on the information returned by each tool, but the agent must repeatedly decide:

  • Which tools are available?
  • Which tool is appropriate for this step?
  • What arguments should I send?
  • Do I have enough information to proceed?
  • Should I retry a failed operation?
  • Should I select a different provider?
  • Is user approval required?
  • Has the overall goal been completed?

This is fundamentally different from traditional applications. Instead of executing workflows that developers have already mapped out. An AI agent determines parts of the workflow dynamically at runtime. That flexibility is what makes agents powerful. It is also what makes them difficult to control.

The LLM Gateway: API Gateway Capabilities for Model Traffic

As large language models became widely adopted, organizations encountered a new set of production concerns. Applications were no longer calling only conventional APIs. They were also sending prompts to multiple model providers, receiving probabilistic outputs, consuming tokens, and incurring variable costs.

This led to the emergence of the LLM Gateway, sometimes called an AI Gateway.

An LLM Gateway applies familiar gateway patterns to model inference traffic. It can centralize capabilities such as:

  • Model and provider routing
  • Fallbacks and retries
  • Token-based rate limits
  • Prompt and response logging
  • Cost and usage tracking
  • Semantic caching
  • Prompt guardrails
  • PII and secret redaction
  • Content filtering
  • Load balancing across models
  • Latency and quality monitoring

For example, an LLM Gateway might route simple classification requests to a smaller, less expensive model while sending complex reasoning tasks to a more capable one. It might switch providers when a model is unavailable, enforce a team’s monthly budget, or redact sensitive information before a prompt leaves the organization.

These are important production capabilities, but the primary object being managed is still the model request.

An LLM Gateway helps an application use models reliably, securely, and cost-effectively.

An agent, however, depends on much more than a model.

The MCP Gateway: Managing Access to Tools and Context

Anthropic introduced the Model Context Protocol, or MCP, in November 2024 as an open standard for connecting AI applications to external tools and data sources. Since then, MCP has become an important part of the agentic ecosystem.

Instead of creating a proprietary integration for every agent and every external system, developers can expose capabilities through MCP servers. An agent can then use those servers to search a database, retrieve a document, create an issue, send a message, query an API, or perform another action.

However, standardizing the protocol does not automatically solve the operational problems involved in managing a large tool ecosystem.

Different MCP servers may represent different security boundaries. They may depend on separate authorization servers, credentials, scopes, network policies, and downstream APIs. One MCP server may use an API key for authentication, while another uses OAuth 2.0, and another OAuth 2.1 + DCR. This makes centrally managing auth difficult. The MCP specification recommends OAuth-based mechanisms, but deployment models and implementation maturity can still vary across servers and a lot of MCP servers today are not fully complaint with the spec.

As organizations connect agents to tens or hundreds of tools, several challenges emerge:

  • How are users and agents authenticated consistently?
  • Which MCP servers should each agent access?
  • Which tools within a server are permitted?
  • Where are credentials stored and exchanged?
  • How can tool calls be audited centrally?
  • How do you revoke access across many servers?
  • How do you prevent sensitive tool results from leaking into prompts?
  • How do you avoid loading hundreds of irrelevant tool definitions into the model’s context?

The final problem is particularly important. Tool definitions and tool results consume context. As the number of available tools grows, blindly exposing all of them to the model can increase token consumption, latency, and the likelihood that the model selects the wrong tool. Tool discovery and selective loading therefore become operational concerns, not merely prompt-engineering concerns.

An MCP Gateway provides a centralized control point between MCP clients and MCP servers.

Depending on the implementation, it can provide:

  • Centralized authentication and credential brokering
  • Authorization and scope enforcement
  • MCP server registration and discovery
  • Tool filtering
  • Tool namespacing
  • Request and response inspection
  • Audit logging
  • Rate limiting
  • Policy enforcement
  • Tool-definition caching
  • On-demand tool discovery
  • Protection against malicious or untrusted tool output

The MCP Gateway governs access to MCP-based capabilities. But MCP servers are still only one part of an agentic system.

The Agent Gateway

An AI agent may depend on some or all of the following:

  • One or more large language models
  • APIs
  • MCP servers and tools
  • Databases and knowledge stores
  • Short-term and long-term memory
  • Other specialized agents
  • Human approval workflows
  • Identity and policy systems

An Agent Gateway is an intelligent middleman between an AI Agent, and all the components it relies on. Agent Gateways provide a centralized control plane and enforcement point across these interactions. These can include, LLMs, MCP Tools, Memory, other Agents, and of course, an API.

Like an API Gateway, an Agent Gateway may proxy requests, enforce access controls, apply policies, limit traffic, and collect telemetry.

Like an LLM Gateway, it may govern model selection, token usage, prompts, responses, safety controls, and cost.

Like an MCP Gateway, it may manage tool discovery, credentials, permissions, and tool invocation.

But its scope is broader than any one of those categories. It governs the agent’s interactions as part of a complete runtime workflow.

An Agent Gateway may be responsible for:

  • Establishing and propagating agent identity
  • Discovering and exposing relevant tools
  • Enforcing policies before and after tool calls
  • Routing requests across models, APIs, tools, and other agents
  • Redacting sensitive data
  • Requiring human approval for high-risk actions
  • Applying budget, token, and execution limits
  • Detecting loops and abnormal behavior
  • Recording end-to-end traces
  • Coordinating access to memory and context
  • Evaluating whether an action is permitted in the current state
  • Terminating an agent run when safety or cost thresholds are exceeded
  • …and many more.

Why Agentic Systems need a Gateway

Agent identity is still an unsolved operational problem

Traditional API security is usually based on a relatively clear identity chain.

A user authenticates to an application. The application receives a token. The token identifies the user, the application, or both. Backend services validate that identity and enforce the corresponding permissions.

Agents complicate this model. An agent may be acting on behalf of a user, running as an independent workload, invoked by another agent, operating across multiple sessions, delegating work to subagents, using shared infrastructure, or even calling tools that require different identities.

When an agent invokes a tool, which identity should the tool evaluate? Is it the identity of the user, the application, the agent, or the organization? What happens when one agent delegates a task to another? Which permissions should be transferred, and for how long?

Without a consistent identity and delegation model, systems tend to fall back to shared API keys or overly broad service accounts. That makes least-privilege authorization difficult and weakens accountability.

An Agent Gateway can provide a consistent point for establishing agent identity, propagating user context, exchanging credentials, narrowing scopes, and recording who or what initiated each action.

Agents expand the security boundary

A traditional chatbot primarily generates text. An agent can generate text and then use that text to take action. It may send an email, modify source code, issue a refund, retrieve customer records, deploy an application, or initiate a payment.

This creates risks that do not exist in an ordinary API traffic:

  • Prompt injection can influence tool selection.
  • Untrusted tool output can manipulate subsequent reasoning.
  • A model can generate valid but unsafe tool arguments.
  • An agent can combine individually harmless operations into a harmful sequence.
  • Credentials may cross boundaries between users, agents, tools, and models.
  • A compromised tool can return instructions disguised as trusted data.

Authentication alone does not solve these problems. A request can be properly authenticated and still be unsafe. Agentic systems need policies that evaluate more than the caller and endpoint. They may need to consider the user’s intent, the selected tool, the arguments, the current workflow state, the sensitivity of the resource, and the potential consequence of the action.

An Agent Gateway can enforce controls at these boundaries. For example, it might allow an agent to search financial transactions but require human approval before issuing a refund.

Agents introduce stateful workflows

An individual LLM request is generally stateless i.e the model receives an input and produces an output.

An agentic workflow is stateful and may maintain conversation history, a task plan, previous tool results, user preferences, intermediate artifacts, approval status, retry counts, budget consumed, and long-term memories.

This does not mean the gateway itself must store all agent memory. In many architectures, memory will remain in dedicated databases, vector stores, or state-management services. The gateway’s role is to govern access to that state.

It can determine which agent may read or write a memory, what context may be sent to a model, which information must be redacted, and whether state from one user or session can be reused in another. This distinction matters. An Agent Gateway does not need to become the memory database.

It needs to become the policy and visibility layer through which memory is accessed.

Agentic observability is difficult

API observability typically focuses on individual requests, such as which endpoint was called, how long it took, what status code was returned, and which service handled the request.

LLM observability adds another set of questions, including which model was used, how many tokens were consumed, what the request cost, and whether a guardrail was triggered. Agentic observability must connect all of these events into a coherent execution trace.

For a single user request, an agent might make several model calls, invoke multiple tools, query memory, delegate to another agent, retry failed operations, and wait for human approval.

When the final result is wrong, it is not enough to know that one API returned a 200 OK; you need to understand the goal the agent received, the plan it formed, the tools available, why it selected a particular tool, which arguments it generated, what the tool returned, how that result affected the next decision, where time and tokens were spent, which policy decisions were applied, which user, agent, and credential were involved, and where the workflow failed.

An Agent Gateway can capture these interactions at a shared boundary and correlate them into an end-to-end trace.

Autonomous execution needs limits

A traditional application usually has bounded execution paths, but an agent can keep planning, calling tools, evaluating results, and retrying until it reaches a stopping condition. Without clear limits, this can lead to long or infinite loops, repeated tool calls, runaway token usage, escalating costs, duplicate side effects, cascading agent-to-agent delegation, and repeated failed authentication attempts.

Conventional request-per-second limits are not enough for agentic systems. They need controls that account for the full workflow, including token and cost budgets, tool-call volume, delegation depth, execution time, retries, concurrency, and the sensitivity of the requested operation.

The Agent Gateway can enforce these limits across the entire execution, rather than evaluating each request in isolation.

From Managing Requests to Governing Intent

API Gateways were designed for a world in which applications knew which endpoints to call.

LLM Gateways emerged when model inference introduced new routing, cost, privacy, and safety requirements.

MCP Gateways emerged as agents gained access to a growing ecosystem of tools and context providers.

Agent Gateways are the next step in that evolution.

They are needed because the unit being governed is no longer only an API request, a model invocation, or a tool call. It is an agentic execution, a sequence of decisions and actions taken dynamically to achieve a goal. That shift alone changes the role of the gateway.

The gateway is no longer concerned only with:

Is this client allowed to call this endpoint?

It must also help answer:

Is this agent allowed to take this action, with this tool, using this identity, in this context, at this point in the workflow?

APIs are not disappearing. They remain the foundation on which agents act, but the clients consuming them are changing. They can reason, choose tools, delegate work, maintain state, and take consequential actions.

As clients evolve, the gateway must evolve with them.

Building the Future: Fabric Gateway

At Postman, we are building what we believe is the next generation of this architecture, we call it the Fabric Gateway.

Fabric Gateway is designed to bring together the capabilities discussed in this article – API governance, model routing, MCP tool management, and agent-level policy enforcement, into a unified control plane for agentic systems.

Our goal is simple: make it possible to safely run agents in production without sacrificing flexibility, speed, or developer experience.

We are currently opening up early access to teams that are building agentic applications and want to help shape this new category.

If you are exploring agents in production, or thinking about how to govern them at scale, we would love to hear from you.

👉 Join the early access and help define what the Agent Gateway should become.

The post Agent Gateway: The Next Evolution of the API Gateway appeared first on Postman Blog.

Read the whole story
alvinashcraft
58 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Agent Skills 101

1 Share

tl;dr
Imagine an AI assistant in your company that can write emails, check expense reports, and run CI pipeline steps — but it has every capability baked into one giant prompt. It’s slow, expensive, brittle, and dangerously able to access systems it shouldn’t. Now imagine the same assistant composed from many small, well-documented, auditable “skills”: a policy-checker skill for expenses, an email-composer skill, a CI-invoker skill. Each skill advertises a compact summary the agent sees up-front and provides richer instructions or safe code only when needed. That’s the difference agent skills make: modularity, safety, and maintainability.

In this post you’ll get:

  • A practical definition of agent skills and why they matter.
  • The SKILL.md packaging convention and how agents discover and use skills.
  • Runnable C# patterns for semantic (prompt) functions and native (C#) functions that agents can call.
  • A small retrieval-augmented (RAG) example (embeddings + local vector store).
  • Concrete operational guidance: progressive disclosure, security, testing, telemetry and CI.

What are agent skills?
A skill is a small, focused package that gives an AI agent a specific capability. A skill typically contains:

  • Metadata (name, summary, inputs/outputs, examples).
  • A description of what the skill is for
  • One or more prompt templates (semantic functions).
  • Optional native code (API wrappers, scripts) that the agent can call.
  • Optional supporting assets (examples, test cases).

Why skills matter

  • Modularity & reuse: Package common workflows once and share them across agents and projects.
  • Progressive disclosure: Agents initially see only compact summaries (low-cost), fetching full prompts or code only when needed (reduced context bloat).
  • Interoperability: Emerging conventions (SKILL.md / skill folders and plugin metadata) let skills be discovered and loaded by different frameworks.
  • Testability & governance: Skills can be versioned, tested, audited, and signed independently from orchestration code.

SKILL.md: the portable skill package
A commonly adopted pattern is a small folder that contains a single SKILL.md file (YAML or structured text) that describes the skill plus optional prompt, code and example files. This single file is how registries, marketplaces and agent orchestrators discover and reason about a skill.

A minimal skill folder:

  • DocumentSummarizer/
  • SKILL.md
  • prompts/
    • summarize.txt
  • examples/
    • example1.txt

Example SKILL.md (YAML-like):
name: Document Summarizer
summary: Summarize documents into a single short paragraph.
description: |
The Document Summarizer skill produces a concise, factual paragraph that preserves
key facts and avoids hallucination. Use it for internal reports and documentation.
inputs:

  • name: input
    type: string
    description: The text to summarize.
    outputs:
  • name: summary
    type: string
    description: One-paragraph summary.
    prompts:
  • file: prompts/summarize.txt
  • prompts/summarize.txt:
  • Summarize the text below in one short paragraph, focusing on main factual points. {{input}}

Two broad kinds of functions inside skills

  • Semantic (prompt) functions: natural language templates packaged as callable functions. These send text to the LLM and return a string output.
  • Native (plugin) functions: compiled code (C#, Python, HTTP endpoints) that expose controlled side-effecting operations (call an API, query a DB, create a ticket).

How agents use skills (discovery → execution)

  1. Discovery: Agent finds available skill folders (local, registry, or marketplace).
  2. Visibility: Agent sees compact metadata (name, summary, I/O).
  3. Planning: Agent decides if a given skill is relevant to the user request.
  4. Progressive disclosure: Agent loads the full prompt or code only on demand.
  5. Execution: Agent calls a semantic function (LLM) or a native function (HTTP or SDK call).
  6. Audit: Invocation logged and reviewed.

C# approach: two patterns

  • Manual skill loader (framework-agnostic, runnable): show how to load SKILL.md, create prompt functions and call an LLM endpoint directly from C#. This is portable and doesn’t require a specific SDK.
  • SDK-based integration (Microsoft Agent Framework): map skills to a Kernel and register C# methods as callable functions. I’ll explain the SDK approach conceptually and how to adapt the manual loader to it.

Runnable sample: a small console app that demonstrates semantic and native skills
This sample is framework-agnostic and uses HTTP calls to a standard LLM API (OpenAI-style). Files provided:

  • AgentSkillsSample.csproj
  • Program.cs
  • SkillLoader.cs
  • Models/SkillDefinition.cs
  • Skills/DocumentSummarizer/SKILL.md
  • Skills/DocumentSummarizer/prompts/summarize.txt
  • Skills/FinanceSkill.cs (native skill)

Project file: AgentSkillsSample.csproj
Exe net7.0 latest

Notes: YamlDotNet reads SKILL.md, Serilog logs audit events, Newtonsoft.Json serializes HTTP payloads. You can substitute other libs (System.Text.Json) if you prefer.

Note: This code was generated by Copilot and has not been fully tested. Caveat Programmer!

using System.Collections.Generic;
public class SkillDefinition
{
   public string Name { get; set; }
   public string Summary { get; set; }
   public string Description { get; set; }
   public List<SkillIO> Inputs { get; set; } = new();
   public List<SkillIO> Outputs { get; set; } = new();
   public List<PromptFile> Prompts { get; set; } = new();
}

public class SkillIO
{
   public string Name { get; set; }
   public string Type { get; set; }
   public string Description { get; set; }
}

public class PromptFile
{
   public string File { get; set; }
}

Skill loader: SkillLoader.cs

using System;
using System.IO;
using YamlDotNet.Serialization;
using YamlDotNet.Serialization.NamingConventions;

public static class SkillLoader
{
// Loads SKILL.md (YAML) and returns SkillDefinition; progressive disclosure: never load prompt files here
public static SkillDefinition LoadSkillMetadata(string skillFolder)
{
var path = Path.Combine(skillFolder, "SKILL.md");
if (!File.Exists(path)) throw new FileNotFoundException("SKILL.md not found", path);   
 var yaml = File.ReadAllText(path);
    var deserializer = new DeserializerBuilder()
        .WithNamingConvention(CamelCaseNamingConvention.Instance)
        .Build();

    var def = deserializer.Deserialize<SkillDefinition>(yaml);
    return def;
}

// Load a prompt file only when invoked
public static string LoadPrompt(string skillFolder, string promptFile)
{
    var path = Path.Combine(skillFolder, promptFile);
    return File.Exists(path) ? File.ReadAllText(path) : throw new FileNotFoundException(promptFile);
}
}

using System;

[AttributeUsage(AttributeTargets.Method)]
public class SkillFunctionAttribute : Attribute
{
   public string Name { get; set; }
   public SkillFunctionAttribute(string name = null) { Name = name; }
}

using System.Threading.Tasks;

public class FinanceSkill
{
[SkillFunction("GetExchangeRate")]
public Task GetExchangeRateAsync(string currency)
{
// Replace with real API call; this is a stub for demo
return Task.FromResult(currency.ToUpper() switch
{
   "EUR" => "1 USD = 0.92 EUR",
   "JPY" => "1 USD = 150 JPY",
   _ => "1 USD = 1 UNKNOWN"
});
}
}

Native skill attribute and example: Skills/SkillFunctionAttribute.cs and Skills/FinanceSkill.cs

using System;
using System.IO;
using System.Linq;
using System.Net.Http;
using System.Text;
using System.Threading.Tasks;
using Newtonsoft.Json;
using Serilog;

class Program
{
static readonly string OPENAI_API_KEY = Environment.GetEnvironmentVariable("OPENAI_API_KEY");
static readonly HttpClient http = new HttpClient();

static async Task Main()
{
    Log.Logger = new LoggerConfiguration()
        .WriteTo.Console()
        .CreateLogger();

    if (string.IsNullOrEmpty(OPENAI_API_KEY))
    {
        Console.WriteLine("Set OPENAI_API_KEY environment variable.");
        return;
    }

    // 1. Discover skills (metadata only)
    var skillFolder = Path.Combine(Directory.GetCurrentDirectory(), "Skills", "DocumentSummarizer");
    var meta = SkillLoader.LoadSkillMetadata(skillFolder);
    Console.WriteLine($"Discovered skill: {meta.Name} - {meta.Summary}");

    // 2. Example: call semantic (prompt) function (progressive disclosure: load prompt on demand)
    var promptTemplate = SkillLoader.LoadPrompt(skillFolder, meta.Prompts.First().File);
    var inputText = File.ReadAllText(Path.Combine(skillFolder, "examples", "example1.txt"));
    var prompt = promptTemplate.Replace("{{input}}", inputText);

    Console.WriteLine("Calling LLM for summary...");
    var summary = await CallChatCompletionAsync(prompt);
    Console.WriteLine("Summary:\n" + summary);

    Log.Information("SkillInvocation: {@skill} {@inputLength}", meta.Name, inputText.Length);

    // 3. Example: call native skill directly via reflection
    var finance = new FinanceSkill();
    var method = typeof(FinanceSkill).GetMethods()
        .FirstOrDefault(m => m.GetCustomAttributes(false).Any(a => a.GetType().Name == "SkillFunctionAttribute"));
    if (method != null)
    {
        var resultTask = (Task<string>)method.Invoke(finance, new object[] { "EUR" });
        var rate = await resultTask;
        Console.WriteLine("Exchange rate (from native skill): " + rate);
        Log.Information("NativeSkillInvocation: {@skill}", "Finance:GetExchangeRate");
    }
}

// Minimal OpenAI Chat Completion HTTP call (simple completion flow). Adjust model as desired.
static async Task<string> CallChatCompletionAsync(string prompt)
{
    var request = new
    {
        model = "gpt-4o-mini",
        messages = new[] { new { role = "user", content = prompt } },
        max_tokens = 400
    };

    var req = new HttpRequestMessage(HttpMethod.Post, "https://api.openai.com/v1/chat/completions");
    req.Headers.Add("Authorization", $"Bearer {OPENAI_API_KEY}");
    req.Content = new StringContent(JsonConvert.SerializeObject(request), Encoding.UTF8, "application/json");

    var resp = await http.SendAsync(req);
    resp.EnsureSuccessStatusCode();
    var json = await resp.Content.ReadAsStringAsync();
    dynamic data = JsonConvert.DeserializeObject<dynamic>(json);
    // Extract the assistant content safely
    var text = (string)data.choices[0].message.content;
    return text.Trim();
}

How this demonstrates key patterns

  • Progressive disclosure: SKILL.md metadata is loaded up-front; prompt files only when the agent invokes a skill.
  • Semantic function: the prompt template is a first-class asset loaded and passed to the LLM.
  • Native function: a C# method annotated with SkillFunctionAttribute is discoverable and callable by the orchestrator (the agent’s planner) via reflection.

RAG + embeddings: a minimal in-memory example
Real RAG pipelines use a vector DB (Pinecone, Weaviate, FAISS) and an embedding endpoint. For a compact demo we’ll:

  • Call an embedding endpoint for documents and queries.
  • Store vectors in memory.
  • Compute cosine similarity for top-k retrieval.
  • Pass retrieved texts into a semantic skill.

Embedding call (OpenAI-style)

static async Task GetEmbeddingAsync(string text)
{
var request = new { model = "text-embedding-3-small", input = text };
var req = new HttpRequestMessage(HttpMethod.Post, "https://api.openai.com/v1/embeddings");
req.Headers.Add("Authorization", $"Bearer {OPENAI_API_KEY}");
req.Content = new StringContent(JsonConvert.SerializeObject(request), Encoding.UTF8, "application/json");
var resp = await http.SendAsync(req);
resp.EnsureSuccessStatusCode();
dynamic data = JsonConvert.DeserializeObject(await resp.Content.ReadAsStringAsync());
var vector = data.data[0].embedding.ToObject();
return vector;
}

Simple in-memory vector store (pseudo code)

  • Store: List<(string id, string text, float[] vector)>
  • Query: compute cosine similarities and return top-k texts.

RAG flow (high level)

  1. Index documents: embed and store vectors.
  2. On user query: embed query, retrieve top-k docs.
  3. Build a RAG context: join top-k texts with separators (truncation/careful length management).
  4. Pass context to the semantic skill prompt and call LLM.

Operational concerns: security, governance, and robustness
Skills are powerful and potentially dangerous. Here are concrete controls and practices.

1) Least privilege for native skills

  • Native skills that perform sensitive actions (modify database, call billing APIs) must run under scoped credentials.
  • Don’t give skills a “root” service principal; instead issue short-lived tokens scoped to exact operations.
  • Example pattern: the orchestrator authenticates the agent using an identity token, then performs an OAuth token-exchange per invocation to obtain least-privilege credentials for the skill runtime.

2) Runtime isolation & sandboxing

  • Execute native plugin code in a sandbox or separate process with tight ACLs.
  • Use containerization (small docker container per plugin) or platform sandboxing (AppDomains are not sufficient) and restrict network access.

3) Audit logging

  • Log: skill name, invocation timestamp, actor (user/session id), input hashes, output hash, status (success/fail), and duration. Redact secrets.
  • Example Serilog event:
    {
    “Timestamp”:”2026-09-10T12:00:00Z”,
    “Event”:”SkillInvocation”,
    “Skill”:”DocumentSummarizer”,
    “User”:”user-123″,
    “InputHash”:”sha256:abcd…”,
    “OutputHash”:”sha256:ef01…”,
    “DurationMs”:312,
    “Result”:”Success”
    }
  • Retention: keep logs for the retention period required by compliance (e.g., 90 days, 1 year). Make them queryable.

4) Approval & registry

  • Keep a registry of approved skills. New skills must pass code review, static analysis, and security review.
  • Require signatures or checksums when installing third-party skills.

5) Limit function-calling from LLM

  • When you advertise native functions to the LLM, provide explicit allow lists. Only expose the signatures the model truly needs.
  • Consider requiring explicit operator approval for risky operations (e.g., “Do you want to run job X? Click approve.”).

Testing & CI for skills
Treat skills like code: add unit tests, integration tests and CI gates.

1) Unit tests

  • Native functions: test logic with xUnit/MSTest, mocking dependencies.
  • Semantic functions: unit test prompt templates by calling a local/mock LLM that returns deterministic output. Mock the HTTP client (HttpMessageHandler) so you can assert prompt content and expected result parsing.

Example xUnit test for FinanceSkill
[Fact]
public async Task GetExchangeRate_ReturnsExpected()
{
var svc = new FinanceSkill();
var res = await svc.GetExchangeRateAsync(“EUR”);
Assert.Contains(“EUR”, res);
}

2) Integration tests

  • A test that runs the orchestration path: index a small set of docs, run RAG + semantic skill, assert returned summary contains a known fact.
  • Use a staging LLM key with usage limits to avoid cost.

3) CI pipeline (GitHub Actions example)
name: CI
on: [push, pull_request]
jobs:
build_and_test:
runs-on: ubuntu-latest
steps:
– uses: actions/checkout@v4
– uses: actions/setup-dotnet@v4
with:
dotnet-version: 7.0.x
– run: dotnet restore
– run: dotnet build –configuration Release –no-restore
– run: dotnet test –no-build –verbosity normal

Add a step for SKILL.md schema validation (use a small script that parses SKILL.md using YamlDotNet and fails CI for missing fields).

Observability & telemetry

  • Track metrics: per-skill invocation count, error rate, latency and model tokens consumed.
  • Export traces: instrument the orchestrator to emit distributed traces (OpenTelemetry) for each skill invocation, so you can correlate LLM calls with native actions in traces.
  • Use Serilog + sink to Application Insights or to a log platform.

Resilience patterns

  • Retries and backoff for failing native calls (use exponential backoff with jitter).
  • Circuit breakers for repeatedly failing or slow skills.
  • Request batching for embedding calls (batch inputs to embedding endpoints to reduce cost).

Advanced patterns and mapping to SDKs

  • Microsoft Agent Framework: these SDKs give built-in primitives that map to this model:
  • Import prompt files as “semantic functions” and native C# methods as plugin functions using attributes like [KernelFunction]/[SKFunction] so the model can call them directly.
  • The SDK advertises function signatures to the LLM (function-calling) and can automatically route decisions.
  • SKILL.md mapping: when loading a skill folder you can:
  • Parse SKILL.md to enumerate functions and prompts.
  • For each prompt, call Kernel.CreateSemanticFunction or your equivalent helper to register it.
  • For native code, reflect over types with attributes and register them as callable functions.

Common failure modes and mitigations
1) Hallucinations / incorrect output

  • Mitigation: RAG (ground responses in retrieved docs), explicit instructions in prompts, ask agent to state uncertainty.
    2) Incorrect skill selection
  • Mitigation: improve skill summaries and include disambiguating examples; add a planner policy that prefers small, well-defined skills.
    3) Malicious or buggy skill
  • Mitigation: vet skills, require signatures, run native skills in sandbox and under restricted credentials.
    4) Escalation loops between skills
  • Mitigation: set depth limits on planner calls, track call stack, and abort when loops detected.
    5) Cost runaway (too many model calls)
  • Mitigation: enforce quotas, use cheaper models for routine tasks, cache results.

When to use semantic vs native skills

  • Semantic skill (prompt) when: the task is pure language transformation or reasoning (summaries, analysis, paraphrase).
  • Native skill when: you need to call external systems, perform deterministic computations, or run actions that must be auditable and auditable (DB writes, sending emails, creating tickets).

Packaging and governance checklist

  • SKILL.md with schema-validated fields (name, summary, inputs, outputs, author, version).
  • Example inputs and expected outputs (test vectors).
  • Unit tests that run in CI.
  • Security review & signing (PGP or similar).
  • Runtime policy: allowed to run? who can install/update?
  • Observability hooks: logging, metrics, tracing.

Key takeaways

  • Agent skills let you compose agents from small, testable, auditable capabilities that are discoverable and loaded on demand.
  • Use SKILL.md-like packaging, progressive disclosure, and clear input/output contracts.
  • For C# developers, you can implement skills yourself (load SKILL.md, create prompt functions, expose native C# methods via reflection or attributes) or use SDKs like Semantic Kernel for structured integration.
  • Secure native code, audit invocations, add tests and CI, instrument telemetry, and enforce governance.

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories