Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162599 stories
·
33 followers

Model Selection Criteria

1 Share

How do you decide what model to use? Let’s look at workload, deployment, regulated environments and other criteria for selecting the right AI model.

In an AI application, a model is the trained component that receives an input and produces generated output such as text or structured data. The models we’re considering are large language models (LLMs), which can interpret language and return text or structured data (e.g., GPT, Claude, Llama, Mistral).

A model can perform well in a prototype and still be the wrong choice for production because the surrounding requirements change what counts as a good fit.

For example, a security review may show that requests cannot cross a regional boundary, while load testing may reveal that the model misses the feature’s response-time budget. Cost at projected traffic and the team’s ability to operate the deployment can also remove models from consideration.

In this article, we’ll look at the practical criteria that shape model selection. We’ll focus on workload fit and deployment requirements, including the constraints that come with regulated environments.

What the Model Needs to Do

Before we compare models, we need a clear description of the job. A strong score on a broad reasoning test won’t tell us whether a model can handle our document formats or return tool arguments that our backend accepts.

An evaluation set built from real inputs gives us a consistent way to compare candidates. Each test needs an expected result or a scoring rule so we can judge every model against the same requirements.

For each model, we can ask four questions:

  • Does it produce useful and accurate output?
  • Can it handle the required context length and languages?
  • Do its structured responses match the schemas our application expects?
  • Do its tool calls select the right operation and provide valid arguments?

A smaller model may be enough for document classification, while an agent choosing tools may need stronger reasoning. Public benchmarks and model cards can help us find candidates before we run these tests. A model card describes intended uses and limitations along with licensing and published evaluations. Those results give us a starting point while our own inputs show whether the model fits the feature we intend to ship.

Where the Model Runs

Once we know a model can do the job, the next question is where it will run. Running a model to produce an output is called inference, and the infrastructure that handles this work becomes part of the selection decision.

With a hosted model API, the provider operates the serving stack. A self-hosted model makes our team responsible for that stack in a cloud account or local environment.

A self-hosted model can run in our cloud account or our own data center. When it runs in our data center, the deployment is on-premises.

There are four common arrangements, each giving us a different level of control and operational responsibility:

DeploymentWhere Inference RunsWho Operates ItStrong FitMain Constraint
Hosted APIProvider infrastructureModel providerFast integration or traffic that changes quicklyRequests follow the provider’s data boundary and service terms
Managed private deploymentAn isolated managed environmentProvider or cloud platformTighter networking or reserved capacityAvailability and model choice vary by provider
Self-hosted cloudOur cloud accountOur teamControl over model versions and serving configurationWe own accelerator capacity and scaling
On-premisesOur data centerOur teamPolicies that require locally operated infrastructureWe own the hardware and its full lifecycle

If we choose to self-host, we need access to the model’s weights (the numerical values learned during training that help determine its output) under a license that allows our intended use. We also need to confirm that the model works with serving software our team can maintain and fits within the GPU capacity available at normal and peak traffic.

Because deployment affects performance and cost, we should compare candidates in the setup we intend to use. A hosted endpoint and a self-hosted deployment of the same model can differ in response time and total cost at production traffic.

Model Selection in Regulated Environments

The deployment choices become more specific when the feature handles regulated data. The applicable rules and our organization’s risk assessment give us concrete requirements for where inference can run and how its data must be handled.

In healthcare, HHS guidance on HIPAA and cloud computing says a covered entity may use a cloud service to process electronic protected health information when the required business associate agreement and HIPAA safeguards are in place. The organization still needs to understand the cloud environment and complete its own risk analysis.

For broker-dealers, FINRA’s cloud guidance says moving infrastructure to the cloud does not remove the firm’s regulatory responsibilities. The deployment still has to support vendor oversight and recordkeeping.

That means we need more than a model that performs well. Before using one with regulated data, we should be able to answer some additional practical questions:

  • Where are prompts and outputs processed, and can the provider use them for training?
  • How is the data retained and deleted, and which access and audit records can we export?
  • Which contracts govern the provider and any subprocessors it uses?
  • How are model changes approved, and can we restore an earlier version?

Making the Final Choice

Once we have a shortlist that fits the workload and deployment requirements, we can make the final choice by answering four questions:

  1. Does the model meet the workload? The model needs to clear the quality threshold on representative inputs and support the capabilities the feature uses.
  2. Can we deploy it within the required boundary? Its license and contracts must permit our intended use, and its data handling has to satisfy the relevant policies.
  3. Can it meet the runtime budget? Response time and total cost need to remain within budget at normal traffic and peak demand.
  4. Can we support the deployment? The responsibilities should be clear, including who manages capacity and model changes.

Together, these questions keep us from choosing a model based on quality alone. Recording the answers also shows why it was selected and gives us a useful starting point when the workload or deployment environment changes.

Wrap-up

Choosing a model begins with the work it must perform and where inference can run. From there, we can compare the remaining candidates against the response-time budget and the cost of the intended deployment.

A hosted model can still be a good fit for a regulated workload when its contracts and controls satisfy the requirements. When inference needs to stay within infrastructure we operate, self-hosting can provide the necessary control. The right model is the one that fits both the feature and the environment in which we need to run it.

For more on building AI-powered applications and agents with Progress, check out the following resources:

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Getting started with DSC 3.0 – Part 6: Using the DSC MCP server with your AI coding tools

1 Share

Throughout this series I ran into the same kind of problem again and again. Which resource type do I need? What is the exact name and casing of a property? Which adapter runs it? Names also change between versions. The Windows PowerShell adapter got a new name in 3.2, and the dsc mcp command became dsc server in 3.3.

An AI assistant that writes DSC YAML for you can easily get this wrong. It guesses, based on what it saw during training. DSC 3.0 has a built-in answer for that: an MCP server that lets the assistant ask DSC itself.

Note: this post is part of a bigger series on Microsoft Desired State Configuration.

DSC has an MCP server built in

You start it with:

dsc server

On DSC 3.2 the command was dsc mcp. From 3.3 on, mcp still works as an alias.

It talks over standard input and output. The AI tool starts the process and communicates with it, so you rarely run it by hand.

Remark: The Model Context Protocol is an open standard that connects AI agents to external tools and data. A server offers tools, and the AI tool, the client, calls them. Your AI coding tool is the client, and DSC is the server.

What tools does the DSC MCP server has to offer?

The tools fall into four groups:

  • Discover: list_dsc_resources lists the resources on the machine, show_dsc_resource shows one resource with its properties, and list_dsc_functions lists the configuration functions
  • Understand: show_dsc_schema (new in 3.3) returns the JSON schema for a configuration document, a resource or an output type
  • Evaluate: invoke_dsc_function and invoke_dsc_expression (both new in 3.3) let the assistant check a function or expression instead of guessing the result
  • Act: invoke_dsc_resource and invoke_dsc_config can run operations on your system. More about those below
Here is the list of tools as seen through the ModelContextInspector:


Why do we need this?

Take Part 4 of this series. We had to get the adapter name right, put requireAdapter on every instance, and use the exact property names of resources like WebSite and WebAppPool. With the server it can look those up on our machine instead of relying on memory.

Notice the last words: our machine. The server only knows the resources and adapters installed where it runs. Therefore, it is important to install the modules first, for example WebAdministrationDsc, or the assistant won't see them.

Set it up

Every client starts the same process. Only the JSON wrapper around it differs.

VS Code with GitHub Copilot

Create .vscode/mcp.json in your workspace:

{
  "servers": {
    "dsc": {
      "type": "stdio",
      "command": "dsc",
      "args": ["server"]
    }
  }
}

Open Copilot Chat in Agent mode, click the tools icon and check that the DSC server shows up in the list.

Claude Code

claude mcp add --scope project dsc -- dsc server

Everything after -- is the command that starts the server. With --scope project the entry is written to .mcp.json in your project root, so you can commit it and share it with your team. That file looks like this:

{
  "mcpServers": {
    "dsc": {
      "command": "dsc",
      "args": ["server"]
    }
  }
}

Run claude mcp list to verify the server is connected.

Try it

Give the agent a concrete task and tell it to use the server.  For example:

Write a DSC v3 configuration document in YAML that installs IIS on Windows Server, creates an application pool, a website on port 8080 and a web application. Use the DSC MCP server to check which resources and adapters are available and what their schemas look like. Use dependsOn between the resources. Don't run anything.

The agent should look up the resources first and then write the document. You can see each tool call in the client. Then do what we did in the earlier parts: run dsc config test and check the result before you apply anything.

Some other things to try:

  • Make sense of an export. Paste the output of dsc config export from Part 3 and ask the assistant to trim it into a desired-state document for the services you care about
  • Check an expression. Ask it to evaluate a configuration expression before you put it in your document
  • Explore. Ask what resources are available for a task, like the documentation's example "How can I manage Windows registry settings?"

Stay in control

An agent that can call invoke_dsc_resource can change your machine. A few habits keep this safe:

  • Mind the permissions. The server is started by your AI agent, so it runs with the permissions of that tool. Use a sandbox or limit the list of available tools.
  • Review the tool calls. Most clients ask for approval before a tool runs. Don't auto-approve the invoke tools.
  • Test first. Use dsc config test and dsc config set --what-if before a real set and try it on a test machine before production. The DSC documentation says this as well: always validate generated configurations in a test environment.
  • Read what it wrote. The assistant is guessing less, but it still guesses. You are responsible for the document that runs.

Wrapping up the series

We started with what DSC is and how 3.0 differs from the earlier versions. Then we managed a service, exported the current state, configured IIS with multiple resources and dependsOn, guarded a configuration with an assertion, and now let an AI assistant ask DSC for the facts.

DSC 3 is a completely different tool than the PowerShell DSC many of us know. Smaller, cross-platform, just data in YAML and in my opinion a lot easier to use.

There are some rough edges and the list of supported resources is still limited (but is growing). I have planned to write an extra post to share some of the issues (and solutions where available) that I encountered. 

I hope at least that this series made the switch a bit easier.

More information

Read the whole story
alvinashcraft
18 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

If one anti-malware software is good, does that make two better?

1 Share

A colleague on the enterprise support team was investigating a complex failure from a customer, and after much analysis, it appeared that the problem was that the system had two different anti-malware programs installed. Let’s call then Contoso and Fabrikam. Fabrikam called a function that Contoso had detoured. Contoso’s detour thought this was suspicious, so it tried to quarantine the Fabrikam process. Unfortunately, Fabrikam had also detoured a function that Contoso was using. The result was mass confusion as the Contoso ended up calling into something that it was trying to quarantine.

Installing two anti-malware programs on the same system is like hiring two different security companies to patrol your building. If they don’t know about each other, each is going to think the other one is an intruder. If you’re lucky, their instructions are merely to report on suspicious activity.

But if you give them weapons and the authority to use them, you may end up with your two security companies pointing their weapons at each other.

In real life, you would introduce the two security companies to each other, or at least make sure they don’t patrol the same floor.

In software, this is harder to do. It’s not like you can invite two programs to lunch and have them get to know each other.

Anti-malware software typically does things that aren’t officially supported, like detouring system functions. And these detours mean that calling a system function no longer does the system thing; instead, it calls into the anti-malware software, and it might decide to do something unrelated to the system function you thought were calling, and that unrelated thing may itself start causing problems, say, by hanging. Not only is there no easy way to identify that this has happened, short of debugging an actual malfunctioning system, but even after you figure out the conflict, there’s usually no way to tell each anti-malware software to “stay on its floor.”

The post If one anti-malware software is good, does that make two better? appeared first on The Old New Thing.

Read the whole story
alvinashcraft
25 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

12 Steps To Self-Editing A Manuscript: A Stress-Free Guide

1 Share

Learn how to self-edit a manuscript with these 12 practical steps, from plot and characters to pacing, continuity, grammar, and final checks.

A Stress-Free Guide To Preparing A Manuscript

When we talk about rewriting in Writers Write, delegates often go a bit pale. They seem to think that you need to rewrite a story or novel from scratch. While certain drafts do need to start their journey again from a blank page, if you have a reasonable first draft, you can follow these 12 chronological steps to make self-editing more manageable.

12 Steps To Self-Editing

1. Read through.

Print out your manuscript, make yourself a coffee, and grab a pencil. Read it from beginning to end as a dispassionate reader. Make the odd comment in the margin if something glaring pops up, but hold back from making detailed notes. The idea is just to get a feel for the global story, the flow, and the feel of it.

2. Plotline.

Now it’s time to interrogate the plot and determine if there’s enough conflict in the story. Look at each scene and sequel to see if you’ve unpacked the major story question posed by the inciting incident. As Sol Stein suggests, compare your strongest scene with your weakest scene. Decide if the weaker one can be recycled or rewritten.

3. Hero in the spotlight.

Here we pick apart the main character. A good idea is to create a character sheet. Then describe their psychological, physical, and socio-economic make-up – read this post. Make sure that every decision or behaviour they display in the story is consistent with these traits.

4. Rattle the cage for the antagonist.

The next step is to do the same for the antagonist. Make sure that they are positioned to bring out the most conflict from your main character. Nothing destroys a story like unfair odds between the hero and their nemesis. Make sure they are equally strong, if not a bit more wily than your main character. If you need to plug more into your plot, go back to step two.

5. Dust off your supporting cast.

To a lesser degree, you will do the same for the other characters in the story. While they may not need the same magnifying glass, you should make sure they’re fulfilling their roles in a vivid, lively and engaging way. A tip is to spend just 20 minutes on each character, freewriting or brainstorming ideas to make them pop. Feed these into the story.

6. Infuse your palette.

Now it’s time to look at setting. Try to put in setting detail where it’s lacking or unclear. Remove places where you’ve been overly descriptive. Make sure you’ve used as many senses as possible to bring these settings to life. Take time out to do research on places you’re unfamiliar with so that these parts of your book hum with authenticity.

7. Talk it out.

If step six asks you to look at the manuscript with a fresh eye, this one demands you bring a keen ear. Read your dialogue aloud or record it and play it back to yourself. Does it sound realistic? It should give us information about the characters – it must tease out their individuality, their background and, at the same time, move the story forward. Read plays or film scripts for inspiration.

8. It’s a sprint, not a marathon.

Now you should look at pacing. Does your manuscript have enough white space? Try to keep sentences and paragraphs as short as possible – just keep in mind that some genres allow for a more leisurely pace than, say, a thriller. If you’re getting bored reading a page, your reader will be too. Be merciless. A tip is to cut every second or third word and see if the story can survive these cuts.

9. Beginnings, middles, and ends.

Look at your first and last page side by side. If you can, try to bring in symbols, images or moods that echo or contrast each other. Find a way to create bookends that will resonate with the reader in a subliminal way.  Now go to the middle of the book – the hinge – and see if this section of the book is a powerful enough mid-point to drive the story towards its climax. It should be a false high point or false low point for the main character and reaffirm his commitment to the story goal.

10. Become a continuity editor.

Put the manuscript away for at least eight weeks, longer if you can manage it. Print out a fresh copy and look for consistency and clarity on every page, every line, in every word. Look for gremlins – a character’s eye-colour changing from one chapter to the next or someone encountering a tiger in Africa. A good way to do this is to imagine each chapter is a stage play – have you signposted your stage directions in a clear, but unobtrusive way.

11. Polish it till it shines.

Now – and only now, we might add – do you do a linear edit of the manuscript. You check spelling, you check grammar, you check that your formatting is consistent. It’s like dressing your book up for a red-carpet event – it needs to be flawless. A sloppy manuscript – no matter how promising – is often passed over for a mediocre story well-presented when it crosses an editor’s desk.

12. Find another eye.

If you have an objective friend, freelance editors, or an online community of beta readers, give them the manuscript to read and encourage constructive feedback. This is the time to put your ego on the backburner and be open-minded. Listen to what they say, take notes and see if their points are valid.

The Last Word

A first draft is where the story begins, not where the work ends. Self-editing gives you the chance to step back, question your choices, strengthen what is working, and fix what is not. Take your time, be ruthless when necessary, and remember that every change should bring the manuscript closer to the book you wanted to write.

After this, it’s time to make final checks and changes and prepare your manuscript for its final journey – to an agent, editor, or printer if you’re self-publishing. Think of your book as your 18-year-old kid going off to college or varsity. You’ve done the best you can, given them warm clothes and a stern lecture – maybe even a flurry of good luck kisses. Now it’s up to your book to stand on its own.

Are you about to edit your book? You may enjoy these useful posts:
  1. An Editing Checklist For Writers
  2. How to rewrite and revise your manuscript
  3. Rewriting- A Checklist For Authors 

The post 12 Steps To Self-Editing A Manuscript: A Stress-Free Guide appeared first on Writers Write.

Read the whole story
alvinashcraft
57 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Preview v0.102.2814.0

1 Share

PowerToys Preview v0.102.2814.0

This is a preview-channel build for the PowerToys 0.102 release train, produced from the stable branch. It follows v0.102.2803.0; the notes below describe changes since that preview, not the complete 0.102 versus 0.101 changelog. This is not the final stable release.

Installer Hashes

Description Filename sha256 hash
Per user - x64 PowerToysUserSetup-0.102.2814.0-x64.exe 0E0BC3040FAFA85373410BB9221F869DCD6B355B1D95E1B56A11E94FC757D170
Per user - ARM64 PowerToysUserSetup-0.102.2814.0-arm64.exe ECA42FE6646ACD8BE81E8BD3C03B46D6AD8CA8F8D04A080280A01209CD4E07A8
Machine wide - x64 PowerToysSetup-0.102.2814.0-x64.exe 1AB750116360130C7A5B9B7DF1BD6795FD4BB3BF9538DADD64892F40D9DCDAD4
Machine wide - ARM64 PowerToysSetup-0.102.2814.0-arm64.exe DD37E0895EE43F306883F08B9ADB5367AFF5087C2A90E67C84D24FF345B72C2E

Highlights

  • Mouse Button Lock: Added opt-in left, right, and middle button locking so dragging and selection can continue without holding a button down.
  • Light Switch: Added a CLI for status, theme switching, and schedule control, with JSON output for automation.
  • Settings: Added a CLI to inspect module states and save enable/disable changes while PowerToys is closed.
  • Command Palette: Fixed dragging packaged apps onto the dock and displaying their loaded icons.
  • Quick Accent: Restored the horizontal scrollbar for overflowing character lists.
  • Shortcut Guide: Improved startup and indexing, including fixes for rapid-reopen crashes and first-activation flicker.
  • Settings: Replaced animated welcome-page GIFs with muted, looping H.264 videos and a fallback image.
  • ZoomIt: Fixed DSC configuration to use actual registry-backed settings and notify the running utility of changes.

Command Palette

  • Fixed dragging packaged apps onto the dock so pinned entries launch correctly and their icons appear when loaded in #51018 by @jiripolasek

Light Switch

  • Added a CLI to query status, set or toggle configured theme targets, and enable or disable schedules on the running Light Switch service, with JSON output for automation in #50459

Mouse Button Lock

  • Added opt-in Mouse Button Lock for left, right, and middle buttons so dragging and selection can continue after a long press without continuously holding a button in #49279 by @MuyuanMS and @owenpkent

Quick Accent

  • Restored the horizontal scrollbar for overflowing character lists so Quick Accent shows the current position without clipping the scrollbar in #50924 by @daverayment

Settings

  • Added a Settings CLI to inspect module states in text or JSON and save enable/disable changes while PowerToys is closed in #50507 by @Copilot and @MuyuanMS
  • Replaced animated welcome-page GIFs with muted, looping H.264 videos, with a PowerToys logo fallback when playback is unavailable in #51057 by @Copilot

Shortcut Guide

  • Improved startup and shortcut indexing, preventing rapid-reopen crashes and initial overlay flicker while picking up PowerToys shortcut changes on the next launch in #50609 by @daverayment

ZoomIt

  • Fixed the ZoomIt DSC resource to read and update actual registry-backed settings, detect configuration drift, and notify a running ZoomIt instance to reload changes in #50551 by @Gijsreyn

Development

  • Removed unused central package entries and documented retained dependency pins without changing the versions resolved by existing projects in #50880 by @Copilot
  • Moved local build logs into artifacts\logs so context-menu packaging no longer fails when trying to include open MSBuild log files in #51056 by @Copilot
  • Removed obsolete analyzer suppressions for retired template code to keep shared build diagnostics easier to maintain in #51061 by @Copilot
  • Promoted IL2081 trimming diagnostics to build errors in the shared AOT configuration while retaining explicit project-level exceptions in #51073 by @Copilot
  • Aligned main's checked-in release train with stable's existing 0.102 setting; the previous preview already used this release train in #51111 by @Copilot and @LegendaryBlair
  • Enabled nullable reference checking across File Explorer preview and thumbnail projects and their unit tests, with annotations and guards that preserve existing preview behavior in #51099 by @Copilot
  • Updated PowerToys Run architecture documentation to point contributors to the current plugin-management and query-processing code in #50519 by @ElenaGe216 and @Gavin-Yau
  • Enabled warnings-as-errors for the shared UI-test framework after resolving its analyzer warnings and documenting narrowly scoped exceptions in #51085 by @Copilot
  • Corrected developer documentation links so readers can reach the referenced Settings, PowerToys Run, and Group Policy files in #51108 by @xThreeh
  • Updated Image Resizer observable models for WinRT and AOT compatibility and made the corresponding MVVM generator diagnostic fail the build while preserving settings serialization in #51067 by @Copilot
  • Corrected native build configuration warnings, removed an unused COM reference, and separated intermediate output folders without changing shipped filenames or output locations in #51090 by @Copilot

Changes needing final review

Two direct integration commits merged the reviewed main changes into stable without separate PRs: 2646ff0204b8 and 1fca166ef506. Both preserve the existing stable-only PowerDisplay diagnostics omission; no new branch-transition removals were identified.

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete

Apple will reportedly debut its first touchscreen MacBook in three weeks

1 Share

Apple is set to introduce a new MacBook Pro with a touchscreen and an updated iPad Mini "on or around" October 27th, Bloomberg reports. If true, that would put the event just two weeks after the October 13th event Apple announced today, which is rumored to focus on smart home products.

Bloomberg says that the late October event will include an "online video presentation" and an "in-person component for the press," which it reports is similar to what Apple is planning for the October 13th event.

At this Mac and iPad event, Bloomberg reports that Apple will announce new MacBook Pros that are lighter than what's available now, have touchscree …

Read the full story at The Verge.

Read the whole story
alvinashcraft
8 hours ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories