Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162073 stories
·
33 followers

The Agentic Data Science Playbook

1 Share

The following article originally appeared on Vanishing Gradients and is being republished here with the authors’ permission

When an AI agent can explore a dataset, choose a modeling approach, run the analysis, and explain its findings, what should the data scientist do?

Traditionally, data scientists chose each step and implemented much of the analysis themselves. Agentic data science changes that division of work: we can delegate an investigation, including methodological choices, while shaping the question, supplying relevant expertise, and challenging the evidence it produces. For AI-native data scientists, choosing the runtime, writing reusable skills, and designing the workflows and feedback that guide the agent are part of the analytical work.

This article provides a playbook for working with data science agents, from setting up an investigation to reviewing its results and carrying lessons into the next assignment. To see why that requires more than a capable model and a business question, consider an experiment we deliberately started with too little guidance. We gave Claude Opus 5.0 a modified version of the public Elliptic dataset and asked, “Build me a model to detect fraudulent nodes.” The dataset is a graph of Bitcoin transactions: each node is a transaction, and an edge represents a flow of bitcoin between transactions. Some transaction nodes carry licit or illicit labels based on the entities that created them; the rest are unlabeled. Each node has a time step indicating when its transaction was broadcast, allowing us to train on earlier time steps and test on later ones. We also wanted separate results for transactions with many connections (high-degree nodes), which mattered most in the intended application. Importantly, we renamed the columns, changed several features, and reindexed the time steps while preserving their order, making the public dataset harder for Claude to recognize.

Claude wrote the code, trained a random forest, and reported an F1 of 0.87 and ROC AUC of 0.99. It had split transactions randomly, mixing earlier and later time steps in both the training and test sets. That test did not measure how the model would perform on transactions from later time steps. Moreover, Claude also used a feature we had planted as a proxy for the fraud label (yes, we tricked it!), giving the model leaked information it would not have when scoring a new transaction. So how do we avoid these situations?

We then supplied the guidance missing from the initial prompt:  We required a temporal holdout, removed the leaking feature, supplied context about how the model would be used in production, and asked for separate reporting on the high-degree nodes that mattered most. Under the corrected evaluation, F1 was 0.70, overall recall was 0.61, and recall on high-degree nodes was 0.21.

The key point is that “Build a fraud detector” left Claude to infer how the model would be used and what would count as success. AI-native data scientists build and direct an analytical process in which agents can investigate, receive feedback, and return evidence for review. The work begins with deciding how much of the investigation to delegate.  Agentic data science is doing data science work with AI agents as teammates. Crucially, the scope of their responsibility can extend well beyond code implementation. An agent can help frame a question, explore data, test a claim, or communicate a result, provided it has the context and tools to do the work, a way to assess its progress and validate its results.

Asking an agent to write a pandas transformation leaves you as the bottleneck, responsible for deciding every next operation. Asking it to investigate a change in customer behavior gives it larger analytical responsibility. It can inspect a result, form another question, choose a method, and continue. The interaction becomes a conversation about the investigation rather than a sequence of requests for code.

The question may be descriptive (what happened?), diagnostic (why did it happen?), predictive (what might happen next?), or prescriptive (what should we do?). The fraud model is predictive; the pricing investigation later in this article is diagnostic and causal. Across these kinds of work, we need to specify the question and intended use, then verify that the evidence supports the answer.

The following five practices are key to agentic data science:

  • Frame the investigation.
  • Equip the agent for the assignment.
  • Organize the work through bounded experiments, competing analyses, or both, according to the question.
  • Review the result independently.
  • Preserve evidence and turn reviewed lessons into reusable expertise.

The first two practices set up the work. The third determines how the investigation proceeds; the fourth tests its claims. Evidence is captured throughout, and the fifth practice carries reviewed lessons into future assignments.

As in agentic software engineering, the agentic data scientist’s two central responsibilities are specification and verification. Specify the question, intended use, and evidence the agent should produce; then verify that its analysis supports the conclusion. Agents can help with both, while the data scientist remains responsible for judging the question and the evidence.

1. Frame the investigation with the agent

Start by discussing the assignment with the agent. Supply the intended use and organizational context, then let it inspect the data and propose an approach. Method selection can be part of its responsibility. Your intervention matters when a proposal changes the question, rests on a questionable assumption, or needs information the agent cannot obtain. Predicting fraud and deciding which flagged entities to investigate, for example, require different evidence about errors and their consequences. A brainstorming skill such as those in Superpowers can help structure that conversation before you turn it into a task prompt.

A useful specification records that shared understanding. It states the decision, relevant constraints, and evidence the investigation should produce. It need not prescribe every step. In the fraud example, “classify transactions from later time steps using only information available when each is scored” matters more than “use a random forest.” The former defines the analytical task, while the latter selects one possible implementation.

You can specify what the investigation must establish without specifying the answer you want. “Determine whether the data support a recommendation” leaves room for an inconclusive result. “Keep trying until you find an effect” does not.

Turn that discussion into a short analytical brief to give the agent as its task prompt. For an assignment like our fraud example, a starting version could read:

TASK PROMPT:
Question: Can we identify fraudulent nodes as they enter the network?
Use: Support investigation, with separate reporting on high-degree nodes.
Available information: Only inputs known when the node is scored.
Agent discretion: Explore data, propose eligible features, choose models.
Return to me: Unclear feature provenance, changes to the target or
population, or a trade-off that requires an operational decision.
Deliverable: Reproducible analysis, temporal evaluation, subgroup errors,
and a recommendation that states what the evidence cannot establish.

Review it with the agent before the investigation proceeds. If exploration reveals that the evidence cannot answer the question, revise the brief explicitly; do not quietly substitute an easier question.

The deliverable may still be a notebook, model, or report prepared outside a production service. You can begin in the workspace where you already do that work.

2. Equip the agent for the assignment

The task prompt tells the agent what to investigate. It also needs to know how the project works, reach the data, run the analysis, and check the result. The harness is the system around the language model that allows this: its tools, runtime, context, permissions, and feedback from its actions. Its runtime is the environment that executes those actions. A language model alone cannot inspect a warehouse, run a simulation, or recover an interrupted statistical model fit. The environment must make those operations possible and return useful evidence about what happened.

Runtime choices are analytical choices as well as engineering choices. Can the agent execute Python or R with the libraries the task needs? Can a long-running fit continue after an interactive session ends? Which scientific libraries should the agent use? Can the agent inspect plots, or does it only see the code that produced them? Can you reproduce the environment in which it reported a result?

An existing agent runtime may provide most of this. Configuring it means deciding what belongs in Markdown, what needs a tool, and what should be checked by a small script. In an investigation like the fraud example, Markdown can hold the brief and data definitions, while a Python script could check that the appropriate temporal validation split is executed. A CSV data extract may be enough for exploration; if the agent needs data warehouse access, a tool exposed through an MCP server can provide it with appropriately scoped, read-only credentials. A sentence in a prompt cannot enforce that access limit.

Take the same care with outputs. Ask the agent to preserve the data reference, code, environment, assumptions, and diagnostics behind its report. A chat transcript is a poor substitute for a runnable analysis. Review becomes much harder when the only surviving artifact is a confident paragraph about what the agent says it did.

Execution is only part of the problem. An agent may know how to fit a model and still misunderstand what the columns mean. It may find five revenue tables and choose the wrong one. A schema rarely explains which customers were eligible for an offer, when a measurement changed, or why the team stopped using an apparently reasonable metric. This is where agent skills and domain knowledge enter. A skill packages instructions and resources for a type of analytical work. It might contain a modeling approach, example code, required diagnostics, and guidance on when to ask for help. Data documentation supplies the organizational meaning: canonical definitions, table grain, known limitations, and the history needed to interpret a result.

A useful skill is specific enough to change the agent’s behavior. “Be rigorous” gives it little to work with. A fraud-modeling skill can require the agent to establish feature availability, evaluate on later observations, and report performance on operationally important subgroups. For example:

For fraud prediction:

Establish what information is available when a node is scored.

Exclude features derived from subsequent investigations or labels.

Fit preprocessing on training data only.

Evaluate on later-arriving nodes and report the required degree groups.

Flag uncertainty about feature provenance before claiming performance.

These instructions leave room to choose a model. They encode reasons that some apparently successful models should be rejected. Where a requirement can be checked reliably in code, the skill can call a script that performs the check and records its result.

Loading every method and every document into every assignment is unnecessary. Give the agent a way to find relevant expertise, including its scope and exceptions. A forecasting skill should not silently impose its evaluation rules on an unrelated retrospective analysis. Nor should a notebook from last year outrank an updated metric definition merely because it offers convenient code to copy.

To put these pieces together locally, begin with a file-and-code agent in a sandboxed project workspace, such as the following:

fraud-investigation/

  AGENTS.md                # Project instructions, where supported by the runtime

  brief.md                 # Agreed question and delegation boundaries

  data-notes.md            # Sources, column meaning, availability times

  skills/fraud.md          # The methodological guidance above

  environment.lock         # Dependency versions, in your tool’s format

  model/                   # Code the investigating agent may change

  results/                 # Experiment log, diagnostics, saved candidates

  review.md                # Acceptance decision and unresolved questions

Use a project instruction file, such as AGENTS.md in runtimes that support it, to explain which context files the agent should read and how to propose updates to them. In other runtimes, provide those instructions through the supported mechanism. Give the sandbox read access to the approved development data and write access to the model and results directories. Keep the final test data outside of the agent’s accessible workspace. The practitioner can run acceptance checks in a separate environment whose evaluator and data the investigating agent cannot modify. A different folder, or version control alone, is not an access boundary.

Now ask the agent to inspect the inputs, identify unresolved questions, and build a baseline. Before allowing repeated experiments, rerun that baseline and examine its feature-availability record, split dates, and subgroup report. This small rehearsal checks whether the setup works all the way from instructions to evidence. A missing subgroup report points to a different problem than a failed package installation. Resolve those problems before giving the agent a longer run.

3. Organize the investigation

With the question framed and the agent equipped, the next choice is how to organize its work. This depends on the intent of the data science problem. For descriptive work, exploratory data analysis may proceed one question and plot at a time. Building a predictive model may support repeated experiments against a fixed evaluator; a causal question may require comparing analyses built on different assumptions.

In a live exploratory analysis on Show Us Your Agent Skills, Eric Ma (Moderna) uses a marimo notebook as a shared workspace with an agent. He explains the protein mutation data, asks for one plot at a time, corrects a color scale that affects interpretation, and chooses the next question from what he sees. The agent edits the notebook and renders the plots; Eric supplies the domain context, checks the artifacts, and owns the interpretation. The reason Eric needed to be in the loop was that human understanding was part of the objective function here!

Use a bounded experiment loop

For predictive modeling, the autoresearcher pattern organizes the work into a repeatable loop: propose a hypothesis, change the model, evaluate it, and keep or revert the change. The agent records each result and uses it to choose the next attempt. Within the scope you give it, it can explore features and model structure as well as parameter values. This is an inner loop within a broader investigation: the data scientist frames the question and sets the evaluation, the agent searches within those boundaries, and the data scientist reviews the result (potentially using an independent agent) before deciding what to do next.

Define what the agent may change, protect the evaluator from those changes, and set a time or compute budget. This makes iteration a bounded task within the investigation.

A compact experiment contract could say:

Improve the supplied baseline within the agreed compute budget.

You may change model code and propose eligible features.

Keep the target definition, validation split, and evaluator fixed.

Record each hypothesis, change, result, and keep-or-revert decision.

Stop at the budget limit or escalate if the evaluation is unsuitable.

Return the best candidate and the experiment history for review.

In a separate exercise with its own baseline and evaluation, we used this pattern to improve a graph neural network trained on the network data from the opening example. We used lower validation loss as the rule for keeping a change; F1 for fraudulent transactions was a separate measure of the resulting classifier. The agent ran 41 experiments while the team slept and retained seven changes that reduced validation loss. On the validation set, loss fell by about 70%, and F1 for fraudulent transactions rose from about 0.72 to 0.82. The log preserved both successful changes and failed attempts, so we could examine how it reached the result.

One candidate had the highest F1 for fraudulent transactions, but the agent rejected it because its validation loss was higher. That followed the selection rule we had set. The experiment history records that choice for the subsequent review.

The autoresearcher pattern works when an objective gives the agent useful feedback on each attempt. But some investigations turn on which assumptions to make, not which candidate scores best. Those tasks need a different way to organize the agent’s work.

Investigate competing explanations

In causal work, no held-out outcome directly reveals what would have happened without an intervention. The agent needs to examine how different analyses construct and test that counterfactual.

In a demonstration from our Master Agentic Data Science course using simulated subscription-business data, we asked: “What did the price increase cost us?” The outcome is daily conversion rate: paid conversions divided by the pool of potential subscribers. Choices about the observation window, counterfactual, exclusions, and validation produce different analytical paths. A final memo usually shows only one.

Two agent runs estimated conversion roughly 16% below their no-price-increase counterfactuals, yet shipped opposing claims.  Run A attributed its estimated drop to a changing pool of potential subscribers and concluded there was “no real effect,” but did not validate that explanation.  Run B backtested its counterfactual, ran a placebo check, and compared six specifications. It reported a robust relative reduction of 15.6%.

The parallel-analysis pattern has independent agents test different choices in the same investigation: one examines the observation window, another tests seasonal assumptions, and we compare their estimates, uncertainty, and diagnostics. Our Decision Lab work extends this approach across analytical paths, using checks to identify unsuitable analyses and unresolved disagreements.

Both approaches give the agent feedback while it works. The fixed evaluator steers the model experiments; diagnostics help it compare causal analyses. The output is a candidate and experiment history, or a set of analyses with their assumptions and checks. Those are the materials for the next task: verifying the claim.

4. Review the result independently

The agentic data scientist now needs to check what that evidence supports, and a fresh agent can help. Give an independent agent reviewer the original brief, data context, code, diagnostics, and final claim. Ask it to reproduce decisive checks and challenge assumptions. In this adversarial review pattern, the agent raises objections it can substantiate; the data scientist judges whether they change the conclusion.

In the fraud exercise, the agent used the same validation data to guide 41 experiments, so the reported gains may partly reflect what worked on that set. Freeze the selected candidate and assess it on an untouched holdout chosen for the intended use, including errors in the groups that matter.

In the pricing exercise, a fresh reviewer challenged Run A’s conclusion. Run A attributed the estimated decline to a changing pool of potential subscribers but provided no evidence for that explanation. The reviewer found that conversion had been rising before the price increase and that placebo interventions in earlier periods did not reproduce the negative effect. Run B’s analysis, which included these validation checks, was selected in the final comparison.

Netflix’s agentic workflow for causal inference implements this division: an actor performs the analysis and diagnostics, while a critic challenges the reasoning and claims. Humans can inspect and rerun the artifacts.

A fresh agent session is not necessarily an independent review if it can read the investigator’s earlier attempts through the workspace or Git history, though! For a check meant to stand on its own, give the reviewer the original brief, final artifact, and data needed for that check, while limiting access to the prior path. The full experiment trail can be examined separately when auditing how the result was reached.

These examples call for different balances of human and agentic verification. In Eric Ma’s EDA, the agent makes plots while Eric checks them and chooses the next question. In the bounded experiment loop, a fixed evaluator checks each candidate before a person reviews the selected model. In the pricing analysis, agentic diagnostics and critique help a data scientist judge what the evidence supports.

The low-low quadrant leaves little basis for trusting a result (in fact, it’s “vibe data science!”). Repeatable checks can move some work toward more agentic verification, while questions that depend on domain understanding or consequential decisions continue to need human judgment.

5. Turn reviewed experience into reusable expertise

The experiment loop and adversarial review both depend on an evidence trail: the saved artifacts that show what the agent did and why a conclusion survived or changed. Preserve that trail throughout each investigation, including data references, code versions, analytical choices, experiment results, diagnostics, and review findings. Keep the failed alternatives and review findings as well as the final report.

This trail has a second use beyond inspecting the current result. Reviewing it with the agent can reveal missing context, recurring mistakes, or methods worth reusing. The next question is which of those lessons should change how the agent approaches a future assignment. Leaving them in a conversation makes that learning difficult to carry forward.

In the fraud example, we deliberately planted a feature that leaked the fraud label. Removing it corrected that analysis. The reusable lesson is to have the agent check proposed features for leakage: where did each feature come from, and would it be available when a new transaction is scored? That requirement can go into a fraud-modeling skill for future investigations.

Before making a lesson into standing guidance, we need to define where it applies. The planted feature was a problem because it carried information unavailable at scoring time, not because it predicted fraud well. The temporal split likewise fits a task involving transactions from later time steps; it is not a rule for every analysis. A skill should capture those conditions so the agent applies the lesson to the right task.

For example:

Lesson: a feature encoded information from the fraud label.

Scope: prospective fraud prediction.

Update: require a documented source and availability time for inputs.

Evaluation: test whether the agent detects outcome-derived inputs

without rejecting legitimate signals merely because they predict well.

This is where evals enter: repeatable tasks with explicit criteria for assessing the data science agent’s behavior. Here, we evaluate how the agent conducts the analysis, not only its model’s predictive performance. The evidence trail supplies concrete failures that can become test cases for proposed changes to its skills or workflow.

Keep the evals, skill versions, and results together. As reviewed assignments reveal new failure modes, expand the cases and rerun them when the agent’s setup changes. The aim is evidence that its analytical behavior improves, rather than a growing collection of instructions that merely sound sensible.

Workflow changes can accumulate in the same way. If a reviewer repeatedly catches a missing diagnostic, move that diagnostic earlier. If a separate reviewer adds cost but never changes the analysis, reconsider its role. If the agent repeatedly asks the same question about a table, improve the data context rather than supplying the answer again in chat.

A completed assignment need not always produce a new skill. A one-off constraint belongs in the project’s notes; a recurring methodological failure may justify standing guidance. That distinction keeps the next investigation from inheriting every exception encountered in the last one.

When other people use the agents you build

When colleagues use an agent without you mediating each request, your local knowledge has to become shared infrastructure. OpenAI’s internal data agent combines institutional context with query evaluations and existing user permissions. Meta’s Analytics Agent draws on prior analytical work and reusable guidance, exposing generated SQL alongside results. Both illustrate why earlier analyses and corrections belong in the system, not only in an analyst’s memory.

In your own work, you can explain an unfamiliar table or catch a misleading conclusion as it appears. When colleagues use the agent directly, that support must be built into the system. Try an assignment with a colleague and note where you need to step in. Missing context belongs in the agent’s guidance; recurring mistakes become evals; questions beyond its remit need a route to a qualified reviewer. Someone must maintain that guidance, and access controls must limit each user’s data access. The analytical principles stay the same, but the agent can no longer depend on you being present for every investigation.

Put the playbook to work

Choose a small investigation you understand well enough to challenge: a model you periodically retrain or a business metric you regularly explain. Give the agent the decision context and ask it to propose an approach. Agree on what it can decide, then let it carry the investigation far enough to produce evidence you can inspect.

At review, pay attention to where your intervention changes the work. Did the agent need a definition only your team knows? Did a diagnostic overturn its conclusion? If that intervention would help on another assignment, make the relevant context or check available there, and test whether it helps.

AI-native data scientists use their expertise to build and direct analytical agents. They turn lessons from reviewing an analysis into skills and checks, then test whether those changes help the agent on future tasks.

The next cohort of our Master Agentic Data Science course starts Oct 6.



Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Organizations need decision-grade knowledge. AI makes it urgent.

1 Share
AI can make the first part remarkably fast. It can find the page, the discussion and the person who might know. The harder work begins when those sources disagree, or when they become stale.
Read the whole story
alvinashcraft
16 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Anyone can start building verified knowledge with Stack Internal

1 Share
Stack Internal transforms your daily work into a living memory that’s shared with the rest of your team. Now anyone can create and share their knowledge in a Stack Internal workspace for free by visiting stackinternal.com.
Read the whole story
alvinashcraft
23 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Getting ready for 2026 results: A look back on Developer Survey findings

1 Share
Read the whole story
alvinashcraft
42 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Skill vs. will: What leaders must solve for AI adoption

1 Share

If you’re leading AI adoption, you may be asking why trained employees aren’t changing how they work. In this post, you’ll learn how to distinguish skill gaps from motivation and system barriers, so you can choose the right response and build an AI-ready workforce.

William Nations is a Senior Learning & Skilling Manager at Microsoft. He is also the creator of The AI-Ready Executive, a thought-leadership video series that helps senior leaders translate AI ambition into practical decisions about their organizations, teams, and business value.

AI adoption requires capability and the conditions to act

Many organizations begin AI adoption by providing access and training. Both are necessary, but neither guarantees that work will change. Training builds individual capability, but whether that capability translates into changed work depends on the organizational conditions that enable employees to apply it.

Research supports this distinction. Microsoft’s 2026 Work Trend Index surveyed 20,000 workers across 10 countries and examined AI’s reported impact on outcomes such as creativity, work quality, collaboration, and the ability to do higher-value work. The research found that culture, manager support, and talent practices account for more than twice the AI impact of individual skill. In other words, individual capability matters, but the conditions surrounding employees can have an even greater influence on whether AI improves their work.

So, given these findings, where should leaders begin? A useful starting point is to examine two factors: whether people have the skills to use AI effectively and the willingness to apply those skills. I call this the “skill versus will” framework. Leaders should also examine whether the surrounding system enables people to act. These are the three key questions:

 

  • Skill: Can people perform the work appropriately and effectively?
  • Will: Do people see the value, feel confident, and intend to change?
  • System: Do workflows, incentives, leadership behavior, governance, and access make the change possible

 

Use skill and will as the first diagnostic lens but examine the surrounding system before concluding that an employee lacks motivation. Before applying the framework, confirm both access and accessibility. Access means employees have the necessary AI tools, licenses, devices, and permissions. Accessibility means those tools, learning experiences, and surrounding workflows are designed so that all employees can use them, including those who use assistive technologies. Both are prerequisites across every part of the framework.

Consider what Microsoft’s 2026 Work Trend Index calls “blocked agency.” Ten percent of workers surveyed had strong AI skills but lacked the organizational conditions to use them. These employees completed the training. They knew the tools. But no one modeled the behavior, the workflow wasn’t redesigned, or experimentation felt unsafe. That is a system barrier that can suppress employees’ willingness or ability to act, not a skill gap.

So, before adding another course, campaign, or mandate, leaders need to understand what is blocking change and choose a response that addresses the actual barrier.

Skill gaps require practice tied to real work

AI fluency involves more than knowing product features or writing prompts. Employees need to recognize when AI is useful, provide relevant context, evaluate outputs, protect sensitive information, and know when human judgment must lead.

Those skills vary by role. A finance leader, seller, developer, and human resources professional may use the same technology, but they make different decisions, work with different information, and manage different risks.

Generic learning can build awareness, but role-based practice is what builds capability.

The strongest learning experiences start with a real problem that teams care about solving. Practicing on meaningful work gives people a reason to build and apply new skills, while creating an opportunity to improve the workflow itself. Teams can compare results with clear quality standards and identify where AI saves time, strengthens thinking, or requires additional review. This approach builds skill and confidence while generating evidence of real business value.

Low motivation often reflects the system, not the employee

When employees hesitate to use AI, leaders may interpret the behavior as resistance. However, hesitation can be a rational response to the work environment.

Employees notice what managers model, what the organization rewards, which mistakes carry consequences, and whether experimentation feels safe. A capable employee may avoid AI if the process requires duplicate work, the approval path is unclear, or the expected benefit is vague. Employees may also have valid questions about quality, data protection, compliance, accountability, and whether AI improves the work.

These concerns should inform the adoption strategy rather than automatically be treated as a lack of enthusiasm. Leaders can strengthen will and remove system barriers by making the purpose practical:

  1. Start with the business problem. Name the priority workflow and define what better work should look like.
  2. Establish clear boundaries for responsible use while preserving appropriate space for experimentation and creativity.
  3. Explain where people remain accountable.
  4. Align workflows, incentives, and recognition with the behavior leaders are asking employees to adopt.

 

In Episode 4 of The AI-Ready Executive, Alan Murray discusses why skepticism is not always a skill deficit and why leaders should understand employees’ concerns before prescribing more training.

Diagnose the barrier before choosing the response

Before diagnosing skill or will, confirm that employees have access to the tools, learning experiences, and workflows they need to participate. Once access is confirmed, leaders can use the framework to identify the primary barrier and select an appropriate response:

 

 

The purpose is to identify what employees need to succeed, so leaders can avoid prescribing more training when the system is the barrier or increasing pressure when people need practice and support.

Five actions to improve AI adoption

  1. Start with the business problem. Name a priority workflow where speed, quality, insight, or customer experience matters, and define what better work should look like.

  2. Model responsible AI use. Show how leaders use AI, evaluate results, and decide when not to automate. A Microsoft study of 1,800 workers found that manager modeling was associated with a 17-point increase in reported AI value and a 30-point increase in trust in agentic AI.1

  3. Set permission and appropriate boundaries. Define where teams can experiment, what information they can use, what review is required, and which decisions remain with people. Provide enough governance to support responsible use without unnecessarily limiting experimentation and creativity.

  4. Align workflows, incentives, and recognition. Recognize teams that remove low-value steps, improve decisions, strengthen customer outcomes, or responsibly spread an effective practice. Only 13 percent of AI users in the 2026 Work Trend Index said they were rewarded for reinventing how they work.2
  5. Measure workflow outcomes. Track whether the work changed, quality improved, time was saved, or effective practices spread. Course completion and tool usage can provide useful signals, but they should not be the final measures of success.

Start here: one priority workflow

Start with one priority workflow. Confirm that employees can participate fully, diagnose whether the primary barrier is skill, will, or the surrounding system, and choose the response that addresses it.

Once the needed capabilities are clear, use AI Skills Navigator to find learning aligned with the skills your team needs to build and apply.

Sources

1. Microsoft. “Research drop: Empowering managers for an AI-first future.” Microsoft Community Hub. View source

2. Microsoft. 2026 Work Trend Index Annual Report: Agents, Human Agency, and Opportunity, p. 29. View report

 

 

 

 

 

 

 

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete

What’s New in Microsoft Copilot | September 2026

1 Share

Welcome to the September 2026 edition of What's New in Microsoft Copilot! Every month, we highlight new features and enhancements to keep Microsoft 365 customers up to date with Copilot features that help your users be more productive and efficient in the apps they use every day.

 

We had some big announcements last week about how we’re introducing the new Copilot to connect the tools people rely on with the next generation of capabilities they’ll need to build, customize and scale AI across work. If you missed it, you can catch up on those announcements here: Introducing the new Copilot with Home, Code, and Autopilot.

Now let’s take a closer look at what else is new this month:

User capabilities: 

IT admin capabilities: 

User capabilities

New design and functionality for Copilot Chat

Copilot experiences in Teams and Outlook get a refreshed look aligned with the broader Copilot app design, simplifying Copilot navigation, streamlining the chat experience, and providing new ways for users to organize work. The updated design appears in both the full-screen Copilot experience in Teams and Outlook and the Copilot side pane. These updates rolled out in September.

 

The upgraded Copilot new tab page in Edge brings search, chat, and web exploration together in a single search box. Copilot-suggested actions and curated work content sit alongside it, so users can move to the next task without hunting for the right entry point. This feature rolls out in October.

Users can now bring agents and skills directly into their prompts for specialized tasks. Users can call up agents or skills inline by typing ‘/’ (agents) or ‘@’ (skills). This enables uninterrupted work and the ability to compose prompts in one continuous flow. This feature rolled out in September.

 

Copilot Chat now better matches search results from scanned PDFs. This enables users to quickly surface relevant information from more of their files, reducing the time spent manually searching through documents. These features rolled out in September.

Copilot in Word, PowerPoint, and Excel on Android brings Copilot-powered editing to mobile devices. Users can now make changes with Copilot wherever they’re working. This creates a more seamless editing experience and makes it easier to stay productive on the go. This feature rolled out earlier this year on iOS, and rolled out on Android in September.

Plugins, connectors, and Work IQ enhancements to the Copilot App

 

New connectors, plugins, and agents help users get more value from Copilot by bringing relevant business data and industry knowledge into their everyday workflows. Additional new connectors, plugins, and agents across industries and functions include:

  • Industrial & Manufacturing: Autodesk, Resilinic, Vista Finance (by Trimble)
  • Professional & Business Services: askpolly
  • Financial Services: Fiscal AI, D&B Risk Analytics, D&B Finance Analytics, Zacks, MT Newswires, PrivCo, S&P Global Kensho, Square, Morningstar Credit Analytics
  • Healthcare & Life Sciences: AMASS, Cortellis Regulatory Intelligence, SciLeads, Causaly, BioRender, ClinicalTrials.gov, PubMed, PopHive, bioRxiv, DrugBank, CAS, Open Targets, NPI Registry, MedlinePlus
  • Energy & Resources: Rystad Energy
  • Human Resources: TechWolf Analyst Agent
  • Software & Technology: Webflow
  • Retail & CPG: NielsenIQ, Cloudinary, Contentsquare

 

The Microsoft plugin registry provides a unified catalog of Microsoft, partner, and custom-built plugins that work across the Copilot app and Microsoft 365 apps. Plugins bring skills, connectors, and other capabilities to Copilot, helping users complete more workflows in the apps they use every day. IT admins can discover and approve plugins once, manage them centrally through the Microsoft 365 admin center and Agent 365, and make them available consistently across Copilot experiences. Developers and partners can publish once, so they can easily introduce trusted solutions at scale while helping IT maintain control. Learn more about plugins and the new plugin registry. This feature rolled out in September.

 

Apps built with SharePoint Framework can now include interactive UI directly in Copilot chat as UX components. This enables users to view information, interact with controls, and take action without leaving the conversation. This feature is rolling out to Frontier in September.

Soon, SharePoint content and action tools will be available in Work IQ, enabling Copilot to use SharePoint content intelligence and take action on that content. This feature is rolling out to Frontier in September.

 

With Record in the Microsoft Copilot mobile app, users can capture in-person conversations and voice notes on the go and turn them into content they can act on. Record brings real-world audio into Copilot and provides an AI-generated transcript and summary in Copilot Chat, so users can easily preserve hallway conversations, customer visits, and ideas that might otherwise get lost. Users can ask follow-up questions, extract action items, or continue working with the content directly in Copilot. This feature rolled out in September.

Teams Phone agent for Copilot in Teams Phone

Copilot can now help customers get answers and the right support faster with Teams Phone Agent, the AI-powered calling receptionist. For customer-facing organizations and departments using Teams Phone, like bank branches or IT Help Desks, the Teams Phone Agent manages calls with repetitive requests, so employees can focus on the conversations that truly need a human touch. With support for over 60 languages, the Teams Phone Agent picks up incoming calls to help customers with answering common questions or booking and rescheduling appointments. Teams Phone Agent can also intelligently route callers to the right department or employee with an AI-generated summary of the conversation, so customers don't have to navigate phone menus or repeat themselves. This feature rolled out in September.

Citations and image modifications for Copilot in Word

Citations in Copilot responses in Word let users see the sources behind Copilot’s answers. Users can view citations to relevant web and work content directly in the response. This provides greater transparency and helps users understand where Copilot’s information comes from. This feature rolled out in September.

 

Now users can arrange and modify images and shapes directly within their Word documents. Copilot makes it easier to create layouts where visual elements look polished and well organized. This helps users create professional-looking documents with less manual effort. This feature rolled out in September.

Skills and connectors for Copilot in PowerPoint

Skills in Copilot for PowerPoint package multi-step workflows into reusable capabilities. Users can call on a Skill to guide Copilot through each step instead of rebuilding the process prompt by prompt. This saves time and helps users produce more cohesive, consistent presentations. Learn more at Brand Kit and Skills in Copilot in PowerPoint.

 

 

Prebuilt Skills in Copilot for PowerPoint help users improve presentations using ready-made capabilities for common tasks. Skills such as Explain this presentation can assess the overall narrative, audience fit, and flow, while Sharpen slide titles can rewrite titles to better summarize each slide. This helps users quickly refine their presentations without having to create detailed prompts or workflows from scratch. These skills rolled out in September.

Administrators can now publish approved Skills to specific users, groups, or the entire organization. This makes an organization’s workflows available directly in PowerPoint, without requiring each person create or upload them individually. This feature rolled out in September.

Brand Kit in PowerPoint can bring together your organization’s visual identity and reusable Skills for common presentation workflows. Brand managers can now include Skills that capture the processes teams use repeatedly alongside guidance for how presentations should look. This helps teams create presentations that are consistent with both the organization’s brand and its ways of working. Learn more at Brand Skill support for Copilot in PowerPoint. This feature rolls out in October.

Connectors in Copilot for PowerPoint let users bring information from the apps and sources they already work with into their presentations. Users can incorporate relevant information from connected sources while building their slides, so they can easily create presentations grounded in the information they use across their work. These connectors rolled out in September.

Copilot in PowerPoint can now continue working on long-running tasks even after users leave PowerPoint, switch devices, go offline, or end their workday. Users can move on to other work without keeping PowerPoint open or monitoring the task, and notifications let them know when Copilot has finished or needs their attention. This gives users greater flexibility to work across desktop, web, and mobile while reducing the time spent waiting for Copilot to complete longer tasks. This feature rolls out in October.

General availability for Copilot in SharePoint

Copilot in SharePoint helps teams create and organize content, ask questions grounded in SharePoint, enrich libraries with AI-generated metadata, and build reusable skills for repeatable business processes. Users can now access advanced AI capabilities directly in SharePoint sites and document libraries, including image generation and editing, AI-generated site analytics reports, and near real-time metadata autofill. Together, these capabilities help organizations organize, understand, and turn content into actionable knowledge while handling more advanced content-management tasks with AI. Copilot in SharePoint rolled out in September.

Copilot Search is becoming the default search experience across SharePoint and OneDrive. It combines natural language understanding, semantic retrieval, and AI-generated answers so people can ask questions naturally and receive grounded answers based on supported content they are authorized to access. This feature is rolling out to Frontier in September.

 

General availability for Copilot in OneDrive

Copilot in OneDrive has been rebuilt on the same foundation as Copilot in SharePoint, creating a more consistent experience for creating content, organizing information, and working with files wherever content lives. Copilot in OneDrive follows the same licensing approach as Copilot in SharePoint. Updated capabilities in Copilot in OneDrive rolled out in September.

IT admin capabilities

Authoritative sources in the Microsoft 365 admin center

Admins can now manage authoritative sources for Copilot Search in the Microsoft 365 admin center. They can designate up to 100 SharePoint sites as authoritative, signaling that the sites contain verified, official content and helping improve the relevance and trustworthiness of Copilot Search and Chat results. From this centralized governance layer, admins can view, add, remove, and bulk import or export authoritative sources. This feature rolled out in September.

Targeted surveys in the Copilot Dashboard

Leaders can now launch targeted Pulse surveys using dynamic survey lists to understand Copilot adoption and impact. The survey audiences are dynamically created from Copilot adoption data and can be filtered using existing criteria, eliminating the need to manually build and maintain lists. By bringing employee sentiment together with Copilot adoption insights, organizations can better understand how users are experiencing AI and identify opportunities to improve adoption and impact. Privacy is maintained through aggregated, population-level reporting rather than individual user visibility. This feature rolled out in September.

 

 

Did you know? The AI at Work Roadmap is where you can get the latest updates on productivity apps and intelligent cloud services. Microsoft Copilot release notes is where you can see the Microsoft Copilot features that are generally available and specific to each platform. Check back regularly to see what features are in development, coming soon and generally available. Please note that the dates mentioned in this article are tentative and subject to change.

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories