Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
158495 stories
·
33 followers

It’s time to panic about AI safety

1 Share

When the phrase "OpenAI hacked Hugging Face" has more or less entered mainstream culture, you know we have an AI problem. This week, we learned more about exactly how OpenAI's agent broke out of a sandbox and autonomously traversed the web, including a bunch of other supposedly secure web services, all in the name of cheating on a benchmark tests.

The fact that this hack happened is a problem. So is the fact that it took a while for anyone to notice. And the fact that it seems no one is willing or able to do much to stop it. (And lest you think it's just an OpenAI problem, since we recorded this episode Anthropic acknowledged its models ha …

Read the full story at The Verge.

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Dispatches from O'Reilly: The best risk mitigation strategy in data? A single source of truth​​​​‌‍​‍​‍‌‍‌​‍‌‍‍‌‌‍‌‌‍‍‌‌‍‍​‍​‍​‍‍​‍​‍‌​‌‍​‌‌‍‍‌‍‍‌‌‌​‌‍‌​‍‍‌‍‍‌‌‍​‍​‍​‍​​‍​‍‌‍‍​‌​‍‌‍‌‌‌‍‌‍​‍​‍​‍‍​‍​‍‌‍‍​‌‌​‌‌​‌​​‌​​‍‍​‍​‍‌‍​‌‍‌‌​​‍‍‌​‌‌​‌‍​‌‌‍​‌‍‍‌‍‌‌‍‌‍‌‌‌​‍‌‍‌‍‌‍​‌‍‌‌​‍‍‌‍​‌‍​‍‌‍‍‌‌‍‍‌‌​‌‍‌‌‌‍‍‌‌​​‍‌‍‌‌‌‍‌​‌‍‍‌‌‌​​‍‌‍‌‌‍‌‍‌​‌‍‌‌​‌‌​​‌​‍‌‍‌‌‌​‌‍‌‌‌‍‍‌‌​‌‍​‌‌‌​‌‍‍‌‌‍‌‍‍​‍‌‍‍‌‌‍‌​​‌​‍​‌‍‌‌​​​​​‌‍​‍​‌‍‌‍‌‌‌‍​‍​‍‌​‌‌​‌​‌‍‌​​​​​‍‌​‌​​‌​‍‌‌‍‌‌​‍‌‌‍​‍‌‍​‌‌‍​‍​​​‍‌​‌​​‌​​​‍‌‌‍​‌​‍‌‌‍​‍​​​‌‌‍‌​‌‍​​‌‍​‍‌‌​‌‍‌‌​​‌‍‌‌​‌‌‍​‍‌‍​‌‍‌‍‌‌‌​​‌‍‌​‌‌​​‍‌​​‌‍​‌‌‌​‌‍‍​​‌‌‌​‌‍‍‌‌‌​‌‍​‌‍‌‌​‌‍​‍‌‍​‌‌​‌‍‌‌‌‌‌‌‌​‍‌‍​​‌‌‍‍​‌‌​‌‌​‌​​‌​​‍‌‌​​‌​​‌​‍‌‌​​‍‌​‌‍​‍‌‌​​‍‌​‌‍‌‍​‌‍‌‌​​‍‍‌​‌‌​‌‍​‌‌‍​‌‍‍‌‍‌‌‍‌‍‌‌‌​‍‌‍‌‍‌‍​‌‍‌‌​‍‍‌‍​‌‍​‍‌‍‌‍‍‌‌‍‌​​‌​‍​‌‍‌‌​​​​​‌‍​‍​‌‍‌‍‌‌‌‍​‍​‍‌​‌‌​‌​‌‍‌​​​​​‍‌​‌​​‌​‍‌‌‍‌‌​‍‌‌‍​‍‌‍​‌‌‍​‍​​​‍‌​‌​​‌​​​‍‌‌‍​‌​‍‌‌‍​‍​​

1 Share
Your semantic layer is a risk mitigation strategy. Not risk in the abstract, compliance-framework sense, but the practical, operational risk that quietly drains organizations every day..​​​​‌‍​‍​‍‌‍‌​‍‌‍‍‌‌‍‌‌‍‍‌‌‍‍​‍​‍​‍‍​‍​‍‌​‌‍​‌‌‍‍‌‍‍‌‌‌​‌‍‌​‍‍‌‍‍‌‌‍​‍​‍​‍​​‍​‍‌‍‍​‌​‍‌‍‌‌‌‍‌‍​‍​‍​‍‍​‍​‍‌‍‍​‌‌​‌‌​‌​​‌​​‍‍​‍​‍‌‍​‌‍‌‌​​‍‍‌​‌‌​‌‍​‌‌‍​‌‍‍‌‍‌‌‍‌‍‌‌‌​‍‌‍‌‍‌‍​‌‍‌‌​‍‍‌‍​‌‍​‍‌‍‍‌‌‍‍‌‌​‌‍‌‌‌‍‍‌‌​​‍‌‍‌‌‌‍‌​‌‍‍‌‌‌​​‍‌‍‌‌‍‌‍‌​‌‍‌‌​‌‌​​‌​‍‌‍‌‌‌​‌‍‌‌‌‍‍‌‌​‌‍​‌‌‌​‌‍‍‌‌‍‌‍‍​‍‌‍‍‌‌‍‌​​‌​‍​‌‍‌‌​​​​​‌‍​‍​‌‍‌‍‌‌‌‍​‍​‍‌​‌‌​‌​‌‍‌​​​​​‍‌​‌​​‌​‍‌‌‍‌‌​‍‌‌‍​‍‌‍​‌‌‍​‍​​​‍‌​‌​​‌​​​‍‌‌‍​‌​‍‌‌‍​‍​​​‌‌‍‌​‌‍​​‌‍​‍‌‌​‌‍‌‌​​‌‍‌‌​‌‌‍​‍‌‍​‌‍‌‍‌‌‌​​‌‍‌​‌‌​​‍‌​​‌‍​‌‌‌​‌‍‍​​‌‌‍‌‌‌‍​‌‍​‌‍‌‌‌​‍‌​​‌‌​​‌‍​‍‌‍​‌‌​‌‍‌‌‌‌‌‌‌​‍‌‍​​‌‌‍‍​‌‌​‌‌​‌​​‌​​‍‌‌​​‌​​‌​‍‌‌​​‍‌​‌‍​‍‌‌​​‍‌​‌‍‌‍​‌‍‌‌​​‍‍‌​‌‌​‌‍​‌‌‍​‌‍‍‌‍‌‌‍‌‍‌‌‌​‍‌‍‌‍‌‍​‌‍‌‌​‍‍‌‍​‌‍​‍‌‍‌‍‍‌‌‍‌​​‌​‍​‌‍‌‌​​​​​‌‍​‍​‌‍‌‍‌‌‌‍​‍​‍‌​‌‌​‌​‌‍‌​​​​​‍‌​‌​​‌​‍‌‌‍‌‌​‍‌‌‍​‍‌‍​‌‌‍​‍​​​‍‌​‌​​‌​​​‍‌‌‍​‌​‍‌‌‍​‍​​​‌‌‍‌​‌‍​​‌‍​‍‌‍‌‌​‌‍‌‌​​‌‍‌‌​‌‌‍​‍‌‍​‌‍‌‍‌‌‌​​‌‍‌​‌‌​​‍‌‍‌​​‌‍​‌‌‌​‌‍‍​​‌‌‍‌‌‌‍​‌‍​‌‍‌‌‌​‍‌​​‌‌​​‍‌‍‌​​‌‍‌‌‌​‍‌​‌​​‌‍‌‌‌‍​‌‌​‌‍‍‌‌‌‍‌‍‌‌​‌‌​​‌‌‌‌‍​‍‌‍​‌‍‍‌‌​‌‍‍​‌‍‌‌‌‍‌​​‍​‍‌‌
Read the whole story
alvinashcraft
33 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

How to govern agentic AI, MCPs, and AI code assistants

1 Share

AI code completion built human review into the process by design. A developer types, a suggestion appears, and a human decides whether to accept it. A person looked at every line before it shipped.

Agentic AI breaks that review loop. An agent can open a merge request, call a tool, modify a CI/CD configuration, and push a change, sometimes without a person reviewing each individual step. Add in the Model Context Protocol (MCP), which lets agents connect to external tools and data sources on their own, and the question engineering leaders are asking shifts. It is no longer "which model writes the best code?" It is "what is this agent allowed to do, and how can we prove what it did?"

GitLab's own research across more than 1,500 developers and technology leaders found that 73% of respondents are concerned about the long-term maintainability of code and 86% agree that without clear governance, AI-generated code can compound technical debt faster than traditional development practice.

In this article, you'll learn how to proactively address some of the challenges organizations are starting to see with governing agentic AI. We'll introduce you to a governance framework for agentic AI in software development and explore what to control, where human review still belongs, how to measure a rollout, and a practical checklist for teams standardizing on GitLab Duo Agent Platform.

Why agentic AI needs a different approach to governance

The security model for an interactive coding assistant is straightforward because a human is in the loop at every step: a developer asks a question, reviews a suggestion, and accepts or rejects it.

Agentic AI in automated workflows requires a different approach to governance. When an agent can run tests, modify configurations, and take multi-step actions across the software delivery lifecycle without a human reviewing each step, the relevant questions change to:

  • What can this agent access?
  • What is it authorized to do?
  • What actions did it take, and can that be proven after the fact?

Most teams are already feeling this challenge. Ninety-two percent of DevSecOps professionals report some governance challenge with AI-generated code, and the specific concerns point directly to the questions above.

The top concerns include:

  • Code attribution: The ability to tell which code was AI-generated versus human-written in the first place.
  • Traceability to intent: Connect AI-generated code back to the business requirement it was meant to satisfy.
  • Documentation that scales: Manual documentation practices don't hold up once agents are generating a growing share of the codebase.

Once an agent is approved for a project, it can typically write, delete, and push changes without someone reviewing the action before it happens. However, you are still accountable for what lands in the codebase regardless of whether a person or an agent made the change.

Governance for agentic AI is not an add-on to code completion governance. It is a different approach, built around identity, permissions, and auditability.

Setting controls for MCPs, agents, model access, and tool permissions

Once agents can call tools and connect to external systems through protocols like MCP, permissioning becomes the control point. A useful governance model answers three questions before any agent runs:

  1. Which agents and flows are allowed?
  2. Where are they allowed to operate?
  3. Which models can the agents use?

In practice, this looks like a few layers working together:

  • A central catalog for agents and flows: Rather than every team standing up its own agent integrations, a shared AI catalog lets administrators decide what gets published, aligned to the organization's existing roles and group structure.
  • Composite identity: Every AI agent's identity should be linked to the human user who requested the action, so activity is never attributable to the agent alone. When an agent tries to access a resource, both principals, the agent and the human who instructed it, need to be authenticated and authorized before access is granted.
  • Tool approval guardrails: Individual agent tools can be set to run autonomously, pause for a human reviewer, or stay blocked outright, so a sensitive action like writing a file or deleting a resource waits for sign-off before it executes.
  • Prompt guardrails: Because agents can process untrusted input (a webpage, an issue comment, a file an attacker controls), the platform needs to detect attempts to hijack agent behavior mid-workflow, not just log it after the fact.

The goal is a control plane that treats agent permissions the same way an organization already treats human permissions: role-based, auditable, and consistent across every project.

Data privacy and self-hosted AI: The questions worth asking

Source code is one of the most sensitive assets an enterprise has, and every AI feature that reads it raises a data-handling question. Before rolling out agentic AI broadly, engineering and security leaders typically want clear answers to a short list of questions:

  • Does the vendor train models on our code?
  • Who owns the inputs and outputs?
  • Where do our subprocessors sit, and can we be notified when that list changes?

For organizations in regulated industries, the answer to "where do our subprocessors sit" often needs to be "nowhere outside our own infrastructure." That is why self-hosting matters as a governance lever, not just a deployment preference. Self-hosted options let a team run its AI agents entirely on infrastructure it controls, while still accurately tracking team usage and satisfying regulators.

Bring-your-own-model support extends that further, letting administrators connect models they have already validated internally and map specific models to certain agent flows. With this method, a sensitive workflow can be pinned to a model the organization trusts while less sensitive workflows use a managed option.

Human-in-the-loop: Decide where review still belongs

Governance does not mean blocking agents from acting autonomously. It means deciding, deliberately, where autonomy ends and review begins. A workable policy usually separates two modes:

  • Interactive work, where a developer is present, sees each suggestion, and approves or rejects it directly. This is the mode most teams already understand from AI code completion.
  • Automated or headless work, such as agents running inside CI/CD pipelines without a developer watching in real time. Here, the review has to happen either before the action (through tool approval guardrails) or immediately after (through an audit trail a human can inspect).

Key decision points worth defining

Code review, testing and validation, and deployment approval should all be defined by answering questions like:

  • Which of these can an agent complete unassisted?
  • Which require a named human sign-off before the change proceeds?

In practice, each checkpoint needs its own enforcement mechanism. This could look like:

  • Merge request approval policies determine who has to sign off before a merge request lands, regardless of whether an agent or a developer opened it.
  • Tool approval guardrails for agents decide, tool by tool, whether an agent's action runs autonomously, pauses for review, or stays blocked.
  • Scanner enforcement holds a change at the pipeline level until it clears the security and quality checks your organization requires.

We recommend creating an organization-wide AI governance policy rather than leaving it to individual team norms. This helps create consistent usage and policies, and helps auditors verify your use of AI.

Added context from the decision layer

There is a longer-term benefit to capturing these approvals as structured records rather than letting them evaporate in a chat thread or a reviewer's memory. An exception granted last quarter, the policy version it was granted under, and who approved it are exactly the kind of organizational judgment that both auditors and future agents need to reference.

Enterprises that treat decisions as durable, queryable events get the added benefit that the next reviewer, human or agent, starts with additional context instead of starting from zero.

Learn more about the decision layer and how capturing decision events can impact your software development.

5 metrics that measure an AI rollout

Five categories are worth tracking from the start, and they should be reviewed together rather than in isolation, since a spike in adoption without a matching look at risk metrics is itself a warning sign.

  1. Adoption: Active users of AI features week over week, number of flows or agents run, and which teams have turned agentic capabilities on versus which have not.
  2. Acceptance and quality: Downstream signals like revert rate on AI-assisted changes and how often AI-authored merge requests pass review without rework. A single acceptance-rate number is a weak proxy for effectiveness on its own, since it says nothing about what happens to the code after it is accepted.
  3. Risk: Track how often tool approval guardrails pause an action for review, how often those pauses result in a blocked or modified action, and whether any agent activity triggers a policy violation.
  4. Remediation: Keep an eye on the scanner coverage across projects, the share of vulnerabilities that get auto-remediated versus manually triaged, and time to resolve findings once flagged.
  5. ROI: Once you have a full quarter of data, explore the ROI of your AI investment. Track how much time developers save on tasks agents now handle and weigh that against the cost per resolved issue or merged change, in credits or compute spent. Check out this tutorial for an in-depth look at your AI ROI — you’ll learn how to transform raw usage data into actionable business insights and ROI calculations.

By tracking all five metrics together, you can catch issues with the AI rollout early. A rollout can look successful on adoption and acceptance while accumulating risk that only shows up in an audit six months later. With the above metrics, you get a holistic view of the value of AI to your organization.

Your GitLab Duo Agent Platform governance checklist

For teams standardizing on GitLab Duo Agent Platform, here is a practical starting checklist before expanding agentic AI beyond a pilot:

  • Review the AI Transparency Center and confirm your understanding of data usage, model vendors, and subprocessor commitments.
  • Decide, at the platform level, which agents and flows are approved for use, and publish them through GitLab’s AI Catalog rather than letting teams configure their own.
  • Set tool approval guardrails as always allow, always ask, or always deny, based on the sensitivity of the tool.
  • Set up composite identity so every agent action is linked to the human who requested it, and access requires both to be authorized.
  • For regulated workloads, evaluate self-hosted deployment and bring-your-own-model options against your data residency requirements.
  • Document explicit human-in-the-loop checkpoints for code review, testing and validation, and deployment approval.
  • Turn on audit event streaming for agent activity so every action lands in the same audit trail your organization already reviews.
  • Define your rollout metrics (adoption, acceptance, risk, remediation, ROI) before expanding past a pilot, and review them together on a recurring cadence.

Revisit the checklist each release cycle. Governance for agentic AI is not a one-time setup. New agents, new tools, and new models often reopen governance questions.

Where agentic AI speed meets enterprise control

Agentic AI changes what needs governing. Code completion asked a developer whether a suggestion was good. Agentic AI asks what an agent is allowed to touch, who approved it, and whether that can be proven later.

Answering those questions fully requires governance built into the platform itself, not layered on afterward. GitLab Duo Agent Platform provides an AI Catalog, approval guardrails, and audit event streaming directly into the platform where the work happens, so governance isn't bolted on after the fact. Teams get AI-assisted speed and enterprise control together, because the guardrails are part of the workflow, not a separate process running alongside it.

Start a free trial of GitLab Duo Agent Platform. On the Free tier, you can sign up in a few simple steps. If you're already on GitLab Premium or Ultimate, you can turn on Duo Agent Platform and use the GitLab Credits included with your subscription.

Read the whole story
alvinashcraft
45 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

What’s New in Advanced Installer 23.9

1 Share
Discover what’s new in Advanced Installer 23.9

Read the whole story
alvinashcraft
51 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

From LLM Theory to Practical Agentic Implementations - Seth Juarez

1 Share
From: NDC
Duration: 59:49
Views: 68

This talk was recorded at NDC Toronto in Toronto, Canada. #ndctoronto #ndcconferences #developer #softwaredeveloper

Attend the next NDC conference near you:
https://ndcconferences.com
https://ndctoronto.com

Subscribe to our YouTube channel and learn every day:
/ @NDC

Follow our Social Media!

https://www.facebook.com/ndcconferences
https://twitter.com/NDC_Conferences
https://www.instagram.com/ndc_conferences/

#ai #llm

Take your understanding of LLMs and agentic primitives from theory all the way to hands-on code. We go behind the scenes of LLM training, loss functions, and evolving output formats. You’ll see how agentic loops, advanced memory architectures, and orchestration techniques transform text generation into interactive systems that act, remember, and automate decisively.

Step through live demos, sample workflows, and open source frameworks, gaining actionable insight into how to safely bridge LLM recommendations with execution. This all-in-one session helps developers design, build, and scale agentic AI solutions--combining technical detail, code samples, and best practices for building robust, orchestrated agentic systems.

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete

Should You Understand Your Entire Python Codebase?

1 Share

Should you understand the entirety of your codebase? How familiar are you with Python’s built-in functions? Christopher Trudeau is back on the show this week with another batch of PyCoder’s Weekly articles and projects.

We discuss a recent article by Sean Goedecke titled “In Defense of Not Understanding Your Codebase.” We dig into the differences in the scale of software projects and the factors that can obscure understanding of an entire codebase. We cover how opinions and practices people often argue for are based on the development practices from decades ago.

Christopher shares his recent video course that explores Python’s built-in functions. The course is divided into sections to help you find the right built-in for tasks involving math, data types, iterables, and I/O. We discuss which of these functions are frequently used in our code.

We also share other articles and projects from the Python community, including recent releases, querying with f-expressions in Django, replacing if-else chains, a Rust-based replacement for Python’s json module, and a couple of Python cheat sheet resources.

This episode is sponsored by HydraDB.

Course Spotlight: Exploring Python’s Built-in Functions

Learn Python’s built-in functions for math, data types, iterables, and I/O, and when to use each to write more Pythonic code.

Topics:

  • 00:00:00 – Introduction
  • 00:02:29 – Christopher’s Python News Song
  • 00:03:18 – Python 3.15.0 Beta 4 Released
  • 00:03:39 – PEP 838: Adding python-version to pyvenv.cfg
  • 00:04:59 – PEP 840: Name Resolution in Class Namespaces
  • 00:06:42 – Stop Using if-else Chains
  • 00:13:46 – Sponsor: HydraDB
  • 00:14:44 – Nifty Django Feature: F Expressions
  • 00:20:23 – In Defense of Not Understanding Your Codebase
  • 00:34:13 – Exploring Python’s Built-in Functions
  • 00:45:26 – Video Course Spotlight
  • 00:46:57 – Python strftime/strptime Directive Cheat Sheet
  • 00:50:22 – Itertools Cheatsheet
  • 00:52:59 – Introducing django-orjson
  • 00:55:33 – Thanks and goodbye

News:

Show Links:

Projects:

Additional Links:

Level up your Python skills with our expert-led courses:

Support the podcast & join our community of Pythonistas





Download audio: https://dts.podtrac.com/redirect.mp3/files.realpython.com/podcasts/RPP_E305_03_PyCoders.750771f2ae67.mp3
Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories