Bill Gates has been reflecting a lot on AI lately, and the process has triggered a stark awakening. Once a staunch AI optimist, the Microsoft cofounder is now deeply pessimistic about what AI means for our collective future.
Having been conspicuously quiet on AI issues recently, Gates is back with a nearly 6,000-word essay seeking to reclaim a central role in shaping the technology globally. Titled "The turbulent AI era is here. The choices we make now are critical," Gates warns the world is not remotely ready for what is coming. Worse still, he notes, "we are not preparing for it."
Gates attempts to chart a path forward, but the essay l …
Ikea has teamed up with Microsoft to launch a new assortment of gaming-inspired furniture and home accessories, coinciding with the 25th anniversary of the first Xbox console. The nine-piece Yxstaby collection (good luck pronouncing that) includes a TV stand, laptop stand, side table, lounge chair, and multi-functional toolbox, with standout products like a stool, cushion, decorative platter, and mirrored display box being specifically designed around recognizable Xbox controller motifs.
The Yxstaby range makes it easier to "bring play into a room and tuck everything away afterwards," according to Ikea's press release, with each piece creat …
“100% of my code is written by [insert whichever AI coding agent is popular right now]!”
You’ve probably heard this claim many times this year.
We wanted to find out how many developers have actually fully outsourced code writing to agents and how the share of agent-generated code differs across regions, tech stacks, and seniority levels.
Luckily, our large-scale, globally representative Developer Ecosystem Survey 2026 gave us the perfect opportunity to uncover the trends. In May–July 2026, we asked over 15,000 professional developers worldwide:
“What percentage of the code that you produced last month for work was …?
The answer options were: 0%, 1%–20%, 21%–40%, …, 81%–99%, 100%, and I don’t know.
Here is what we found:
Expectedly, manual coding is disappearing quickly, but not everybody has made the leap to a fully agentic development workflow yet.

On average, professional developers report that:
However, adoption varies wildly across the board. Over half of all developers now write less than 20% of their code manually, and one in five writes literally zero code without AI help. At the same time, the group that relies almost entirely on coding agents (over 80% agent-generated code) remains a minority of around 22%. Most developers are sitting somewhere in the middle.
Interestingly, senior developers are among the first to hand coding over to agents. About a quarter of senior developers generate the vast majority of their code (over 80%) using agents, compared to a smaller fraction of juniors who tend to lean more toward AI-assisted workflows rather than fully agentic coding.
That said, not all seniors are agentic-first yet. Adoption varies widely within this group. See the charts below.

About 32% of developers who report Claude Code as their most-used AI coding tool generate over 80% of their code with agents.
Interestingly, among developers who use Codex, this share is notably higher at 42%. The share of developers who don’t write code without AI assistance at all is 37% among Codex users, which is tangibly higher than among users of other AI coding tools.
Our interpretation is that while Claude Code is increasingly becoming the mainstream AI coding tool (already the de facto standard, with 39% adoption at work), its audience no longer consists predominantly of advanced users of agents. Codex, on the other hand, is catching up in terms of awareness and adoption, and its user base might have more advanced users seeking better value for money (Codex has historically offered higher quotas).
Cursor users are similar to Claude Code users – on average, 58% of their code is agent-generated, and for 28% of its users, coding agents generate over 80% of their code.

There is a clear split in agentic coding adoption by tech stack. Developers with Go, JavaScript, and TypeScript as their main programming languages report the highest shares of agent-generated code, averaging 54%–55%.
On the other end of the spectrum, C and C++ developers remain the least agentic, maintaining a much higher proportion of manually written code – 38% on average.
Java and Python developers sit in the middle with 48%–51%.

The geographic differences are remarkable. Developers across East Asia, particularly in China, Japan, and South Korea, are leading the charge in agentic coding: About twice as many developers there (32%–35%) generate the vast majority of their code (over 80%) with agents, compared with around 16% of developers in Europe and the UK.

We were curious whether developers could be grouped into homogeneous segments based on how they write code today. Three distinct profiles emerged:
See the methodology notes for more details on how we did this.
We deliberately use the word “coders” here to highlight that this is about the code generation process, not development as a whole.
Agentic coders (~31% of developers) mostly write code with agents nowadays:
Despite being the locomotives of agentic coding, only 46%–57% of heavy users of Claude Code and Codex are agentic coders.
AI-assisted coders (~47% of developers) haven’t gone full agentic yet, but already write a large portion of their code with agents. They still prefer an AI-assisted development workflow and don’t hesitate to write some code manually if needed:
Manual coders (~23% of developers) write most of their code manually, though they use AI sometimes, mostly in an AI-assisted rather than fully agentic manner:

Whether you are all-in on agentic workflows or prefer hands-on coding, the shift is clear: Purely manual coding is quickly becoming a thing of the past.
In our previous blog post, based on the same survey data, we explored trends in AI coding agent adoption across the industry and the main players in this market.
We plan to share more materials on agentic development from the Developer Ecosystem Survey 2026 with the community soon.
Stay tuned and subscribe to JetBrains Research blog updates below!
Unrealistic responses – where the sum of the lower bounds of selected answers exceeded 150% or the sum of the upper bounds of selected answers fell below 80% – were not used in this analysis. This filter was applied on top of the regular data-cleaning filters used for Developer Ecosystem Survey data.
We used the midpoint of each answer bucket (0%, 10.5%, 30.5%, 50.5%, 70.5%, 90%, 100%) to calculate averages. The averages across the three categories of how code is written within the same group (e.g. seniors) could exceed 100% because of the bucketed nature of the answers, and respondents’ self-reports may not always be fully accurate.
We employed hierarchical cluster analysis with the Ward method and Euclidean distance on unstandardized bucket midpoints (e.g. 0%, 10.5%, 30.5%) to clusterize developers into homogeneous segments based on how they write code.
In this report, “professional developers” refers to respondents who reported being involved in coding or programming in any of the following job roles:
Roughly 90% of the sample falls into the Developer / Programmer / Software Engineer job category.
The Developer Ecosystem Survey is localized into eight languages: English, Spanish, Chinese, Japanese, Korean, German, French, and Portuguese. We apply quotas on the required number of responses by region to help achieve accurate global representation. The quotas are proportionate to the number of developers in each region, based on estimates by our Data Science team. The detailed methodology of these estimates is described here.
The Developer Ecosystem Survey has been statistically reweighted to better represent the global developer population by region, employment status, programming language, and familiarity with JetBrains products (to avoid data skewed toward an excessively JetBrains-familiar audience). You can read about the weighting methodology for the Developer Ecosystem Survey here.
AI is creating a new layer of enterprise infrastructure. Gateways, retrieval platforms, orchestration services, and containerized runtimes now sit between users, applications, data, and models. These systems concentrate credentials, data access, model connectivity, and execution privileges, making them some of the most powerful components in the AI stack.
That concentration of trust is also creating new opportunities for attackers. In recent investigations, Microsoft observed activity targeting three distinct AI workloads: a LiteLLM gateway, a RAGFlow deployment, and a Kestra workflow environment. The intrusion paths varied, but the objectives were strikingly similar. Attackers sought to steal credentials, establish persistence, and monetize compromised compute resources.
The individual techniques matter, but the broader pattern matters more. Across these cases, attackers treated AI infrastructure as a control plane where credential theft, host compromise, and downstream data access can converge. As organizations continue to deploy AI systems, these platforms are becoming high value targets that deserve the same security scrutiny as other critical enterprise infrastructure.
The campaign-level signal extends beyond one product. The targeted workloads served different functions, but each exposed assets that could support follow-on abuse, including model-provider keys, proxy-issued virtual keys, database connection strings, tenant configuration, workflow execution, or host compute. Post-compromise behavior varied by workload role. Defenders should inventory exposed AI management surfaces, restrict administrative access, and monitor for gateway-originated execution and secret access.
| AI workload | Observed activity | Attacker objective |
| LiteLLM | Observed attacker activity: Python droppers, runtime secret harvesting, PostgreSQL collection, miner deployment, and persistence activity from the LiteLLM gateway context. Microsoft assessment: Initial access likely occurred through exploitation of the exposed LiteLLM gateway surface, consistent with the vulnerability chain involving CVE-2026-42271 and CVE-2026-48710. | Credential theft, backend database access, durable host access, and compute monetization. |
| RAGFlow | Observed attacker activity: Possible SSRF-style reconnaissance followed several days later by code execution, application-path modification, and placement of a Python hook in the TenantLLM credential-configuration flow. Public research: Describes multiple RAGFlow execution paths; Microsoft does not attribute this intrusion to a specific vulnerability. | Intercept newly configured LLM provider credentials and model metadata. |
| Kestra | Observed attacker activity: Workflow-origin shell execution, Docker and container-environment discovery, XMRig deployment, and follow-on data collection. Microsoft assessment: Initial access likely involved exploitation of the exposed Kestra orchestration surface, with CVE-2026-49869 providing relevant public vulnerability context. | Secret discovery, container-level access, data collection, and rapid compute monetization. |
LiteLLM is commonly deployed as a proxy or gateway between applications and model providers. In that position, the service may hold or retrieve model-provider keys, LiteLLM master keys, virtual-key records, database connection strings, routing configuration, and tenant policy data. Command execution in the gateway runtime therefore exposed a process context close to AI routing and credential material.

Initial access
Microsoft assesses with high confidence that initial access likely occurred through exploitation of the exposed LiteLLM gateway surface. Relevant public vulnerability paths include CVE-2026-42271, an authenticated command-execution issue in LiteLLM MCP stdio test endpoints, and the route described in public research that chains this flaw with CVE-2026-48710, a Starlette host-header validation bypass, to achieve unauthenticated remote code execution in vulnerable exposed deployments.
In this chain, CVE-2026-42271 provides the command execution capability through the MCP stdio test path, while CVE-2026-48710 can weaken the authentication boundary in affected configurations, potentially making that capability reachable without valid credentials.
In this case, initial access occurred in the context of the LiteLLM gateway process. The gateway service, rather than an unrelated system process, became the execution origin. Subsequent activity from that point is described in the observed attack chain below.

The first observed stage was credential harvesting from the LiteLLM gateway runtime. The payload read the gateway process environment and filtered for credential-related values, including model-provider API keys, the LiteLLM master key, database connection strings, UI credentials, tokens, passwords, and other secret-like fields.

In containerized LiteLLM deployments where the gateway runs as PID 1, /proc/1/environ exposes the environment block for the gateway process. Telemetry showed the payload reading /proc/1/environ, filtering for keywords such as master, API key, token, password, and UI-related fields, then sending collected values to attacker-controlled infrastructure.
The exfiltration logic used multiple transports in sequence, including Python urllib, curl, and wget. This provided fallback paths if one tool was unavailable or if egress controls affected one outbound method.
The second stage moved from gateway-level command execution to payload delivery. The first delivery path launched from the compromised LiteLLM gateway process as an inline Python command. The code retrieved a masqueraded ELF binary from attacker-controlled infrastructure, staged it under a temporary path, marked it executable, and launched it with command-line arguments resembling a Linux service process.
The downloaded ELF used service-style naming and arguments to masquerade as a benign Linux daemon.
A second delivery path used a shell-stage downloader. A gateway-spawned Python command invoked a shell that used multiple download methods with short timeouts and fallback behavior, staged the retrieved content under randomized temporary paths, marked it executable, and launched it with supplied parameters. Together, these paths show redundant payload retrieval and execution from the gateway process context.

The third stage performed host discovery from the second-stage payload. Observed commands fingerprinted the host, checked privilege boundaries, inspected listening ports, and searched for other miner or remote-access activity on the system.

Relevant artifacts included a silent passwordless-sudo check, listening-port inspection, and multiple process sweeps for miner or remote-access tooling. The searches included terms associated with miners, remote shells, competing infrastructure, and previously dropped scripts.
The fourth stage prepared the host for cryptocurrency mining and removed competing persistence. Telemetry showed XMRig or XMRig-like mining components, RandomX-related tuning, and cleanup commands targeting other miners.

Two artifacts were central. A command loaded the Linux Model-Specific Register module with write access enabled, a behavior commonly associated with RandomX/XMRig CPU tuning. Separately, a crontab rewrite removed entries associated with other miner names before writing the modified cron data back.
The fifth stage used the previously collected database connection string to access the LiteLLM-backed PostgreSQL data tier. This moved the activity from host-level secret exposure into application-layer data collection from the gateway backend database.


Telemetry captured a self-contained python3 one-liner that read the gateway PID 1 environment, parsed DATABASE_URL, installed or imported PostgreSQL support, connected to the database, and dumped LiteLLM model and virtual-key records. The connection string referenced <name>.postgres.database.azure.com, indicating Azure Database for PostgreSQL.
The payload targeted LiteLLM tables including LiteLLM_ProxyModelTable and LiteLLM_VerificationToken. The collected records could include model configuration, upstream provider key material, provider endpoints, and proxy-issued virtual keys.
Output was base64-encoded and exfiltrated in small chunks to an out-of-band callback endpoint. A sibling variant posted data to a separate web endpoint that was also observed during the earlier credential-harvesting stage.
The sixth stage added persistence, command-and-control, and defence-evasion mechanisms. Observed artifacts included service-account SSH authorized-key modification, hidden-file relay execution, masqueraded service names, self-relaunch loops, and immutable-file attributes.

The durable access artifact was an authorized_keys write under a service account. Additional artifacts included hidden-file relay execution, command-and-control relay components, masqueraded systemd service names, and relaunch paths under hidden temporary files.
Names used in relaunch paths overlapped with common Linux daemon naming patterns. Periodic out-of-band callbacks were also observed, providing network telemetry that the payload continued to execute and retained outbound connectivity.
The LiteLLM compromise produced multiple impact paths: provider credential exposure, proxy-issued key exposure, database-backed configuration access, host resource abuse, and durable service-account access. The gateway role made these impacts broader than a standard single-process application compromise.
RAGFlow supports document-processing and retrieval-augmented generation workflows and stores tenant LLM configuration. The observed execution occurred inside the RAGFlow container under the application runtime lineage. That context is important because the affected code paths process provider credentials when users add or modify LLM settings.

Microsoft assesses with high confidence that initial access likely occurred through exploitation of the exposed RAGFlow application surface. Telemetry showed the RAGFlow server process retrieving an attacker-supplied URL through the application’s own HTTP client, resulting in an outbound Burp Collaborator callback without corresponding child-process execution. Remote code execution in the same service context followed later in the observed sequence.
Microsoft assesses with low confidence which specific vulnerability, if any, enabled that code execution. Because the relevant application code paths execute within the RAGFlow Flask service process, endpoint telemetry could not distinguish the precise execution sink. Publicly documented vulnerabilities affecting relevant RAGFlow versions include CVE-2026-45312 and CVE-2026-28797, authenticated Jinja2 server-side template injection issues in the prompt generator and Agent workflow components; CVE-2026-24770, a MinerU parser path-traversal issue that can permit arbitrary file overwrite and subsequent code execution; and CVE-2025-68700, a Canvas CodeExec sandbox-bypass issue tracked as GHSA-8xw3-v6c2-j84j.
These vulnerabilities provide plausible technical context but are not attributed as the confirmed cause of this intrusion. Depending on the affected version and deployment configuration, access to authenticated functionality could also be influenced by separate account-access weaknesses, including CVE-2025-69286. For defenders, the possible SSRF activity through the OASTify relay network is a useful precursor signal because remote code execution in the same service context followed several days later.
The first payload stage located the RAGFlow installation from inside the container and identified the tenant LLM model-configuration path. Telemetry showed discovery logic for common application locations, followed by creation of a hidden runtime hook under the application tree.

The second stage modified the application startup or import path so the hidden hook would load with the RAGFlow service. This tied the credential-interception behavior to the application runtime rather than to a separate long-running process.

The hook wrapped the tenant LLM configuration flow and captured newly supplied provider metadata during credential setup. Captured fields included provider type, model name, API key material, and related endpoint metadata. The collection routine used outbound HTTP from within the container and suppressed errors so the application flow could continue if collection failed.

The final stage wrote or refreshed the hook and created a local marker indicating that installation had completed. Command-line telemetry was partially truncated, but the repeated execution sequence, process lineage, and application-file modifications were sufficient to reconstruct the functional behavior.

The RAGFlow compromise was primarily focused on LLM credential collection rather than host monetization. Telemetry did not show miner deployment or an interactive reverse shell in this case. The affected runtime path could capture provider credentials configured after the hook was installed, and the startup-path modification could persist across service restarts if the modified filesystem state remained present. SSH-key material was also written inside the container, but its durability depends on container privileges, filesystem persistence, and host-container boundary configuration.
Kestra is a workflow orchestration environment. Because workflows are designed to execute tasks and interact with external systems, abuse of workflow-creation and execution capabilities can provide direct code execution in the worker runtime.

Microsoft assesses with high confidence that initial access likely occurred through exploitation of CVE-2026-49869, a critical authentication-bypass vulnerability in Kestra. Exploitation could allow an unauthenticated remote attacker with network access to bypass the login mechanism, define a malicious workflow using the Process runner, and trigger worker-side shell-script execution.
Following the assessed initial-access sequence, telemetry showed two closely timed workflow-origin shell sessions. The first produced shell initialization activity, while the second performed the main follow-on actions, including Docker socket access, container-environment enumeration, miner deployment, and defence-evasion file operations. A later workflow-origin event used a curl-pipe-shell delivery pattern to retrieve remote script content directly into a shell and store collected output through the application’s own key-value interface.
Telemetry showed the Kestra worker lineage spawning shell activity from the orchestration layer. Two closely timed workflow-origin shell sessions were observed; the first produced shell initialization activity, while the second performed the main follow-on actions.
After workflow-origin execution, commands accessed the mounted Docker socket from inside the compromised orchestration environment. The activity queried container metadata and inspected container environment arrays, exposing environment-backed values from other containers reachable through the mounted runtime socket.
This behavior is significant because workflow engines often run near automation secrets. Environment arrays, mounted configuration, service credentials, and container metadata may expose cloud keys, database passwords, API tokens, or internal service endpoints when the container runtime socket is accessible.

The monetization phase followed the workflow-origin execution chain. Telemetry showed miner retrieval from a public release source, archive extraction, binary renaming, background execution, and mining-pool communication. CPU-tuning behavior commonly associated with RandomX/XMRig mining was also observed.
Additional defence-evasion file operations were observed around a temporary path, including restrictive permissions and immutable-file attributes. These artifacts provide file-system telemetry alongside the workflow-origin process lineage and network activity.

A later workflow-origin event used a curl-pipe-shell pattern for follow-on collection. Remote script content was retrieved and executed directly by the shell without being written as a standalone script file first. The resulting output was encoded and stored through Kestra’s own key-value interface.

The Kestra compromise exposed four impact paths: shell execution through the workflow engine, container-environment exposure through Docker socket access, host resource hijacking through miner deployment, and follow-on collection through workflow task execution. The later curl-pipe-shell event encoded collected output and stored it through Kestra’s own key-value interface, reducing reliance on standalone file artifacts.
Several payloads exhibited characteristics often associated with assisted or generated code, including organized imports, explicit timeout handling, dependency fallbacks, formatted output, defensive exception handling, and explanatory comments. Compared with minimal, one-off shell payloads, these samples showed a more structured and robust implementation style.


These characteristics are observations about the tooling, not evidence of attribution. From a security perspective, their significance is that they can improve payload portability and resilience across Linux and container environments. No conclusion about the code’s authorship or development method is required.
Initial access differed by workload. LiteLLM involved command execution from the gateway runtime. RAGFlow progressed from SSRF-style probing to runtime modification. Kestra used workflow execution as the shell-access path.
The observed objectives were consistent. Across the cases, telemetry showed credential collection, durable access mechanisms, and resource monetization, even though the execution path differed by product.
Payload behavior was specific to each workload. LiteLLM payloads targeted gateway environment variables and database-backed proxy records. RAGFlow activity targeted LLM credential configuration. Kestra activity focused on workflow execution, container discovery, and cryptomining.
What this means for defenders: Defenders should monitor AI workloads according to their control-plane role, not only as isolated applications. Gateway, retrieval, and orchestration services can concentrate credentials, database access, workflow execution, and container privileges in one runtime. High-value detections should therefore correlate unexpected application-origin shells or interpreters with secret access, application-file modification, Docker socket use, outbound callbacks, and resource-hijacking activity. Treating these signals as a connected compromise path can expose attacks earlier than product-specific indicators alone.
Microsoft recommends the following mitigations to help reduce the risk and impact of AI workload compromise.
Enable Microsoft Defender for Endpoint protections on Linux. Keep real-time and cloud-delivered protection enabled to detect files written to disk, newly observed droppers, miners, and second-stage payloads. Enable behavior monitoring for anomalous child processes, credential access, data staging, exfiltration, and persistence activity.
Microsoft Defender coordinates detection, prevention, investigation, and response across endpoints, identities, cloud workloads, and apps to provide integrated protection against attacks on AI infrastructure like the one discussed in this blog. Given the criticality of this new attack layer, defender is providing differentiated visibility, detection and protection from attacks against AI resources. Customers with provisioned access can also use Microsoft Security Copilot in Microsoft Defender to investigate and respond to incidents, hunt for threats, and protect their organization with relevant threat intelligence.
| Tactic | Observed activity | Microsoft Defender coverage |
| Initial Access | Exploitation of internet-exposed AI workload surfaces, including model gateway, retrieval, and workflow orchestration services reachable without network restriction. | Microsoft Defender for Endpoint – Suspicious shell execution from an AI workload process – Suspicious shell execution from a scripting application runtime |
| Credential Access | LiteLLM: reads of /proc/1/environ and the model-config table to harvest provider API keys and the database connection string. RAGFlow: TenantLLM.insert() monkey-patched to intercept provider API keys (OpenAI, Azure, Anthropic, Gemini) on every LLM configuration event, exfiltrated to a secondary C2 endpoint. Kestra: Docker socket used to enumerate container Config.Env arrays across all running containers, collecting embedded cloud, database, and API secrets. | Microsoft Defender for Endpoint – Suspicious process collected data from local system – Suspicious file copy operations Enumeration of files with sensitive data |
| Execution & Defense Evasion | LiteLLM: second-stage binary dropped to /tmp and executed under names impersonating system services and daemons. RAGFlow: base64-encoded Python payloads decoded and written to /tmp, executed sequentially to discover the RAGFlow install, inject a persistence hook, and verify implant success — fully automated with no interactive shell. Kestra: malicious workflow submitted via the pipeline API caused the Java worker to spawn a bash reverse shell; XMRig was downloaded, unpacked, and renamed to evade name-based detection. | Microsoft Defender for Endpoint – Hidden file executed – Suspicious process launched from a world-writable directory – Suspicious path deletion – Suspicious file dropped and launched – Suspicious shell command execution – Suspicious piped command launched – Executable permission added to file or directory Possible reverse shell – Suspicious Python command-line execution\ – Suspicious script launched – Process launched in the background – Suspicious file or information obfuscation detected – Suspicious deletion of launched process binary – Suspicious shell execution from a scripting application runtime |
| Impact (Resource Hijacking) | LiteLLM: trojanized runtime binary maintained persistent access enabling ongoing credential and compute abuse. RAGFlow: every LLM API key configured after infection silently exfiltrated, enabling unauthorized use of provider accounts at the attacker’s direction. Kestra: XMRig v6.26.0 launched with RandomX MSR tuning toward a Monero mining pool, consuming host CPU for attacker profit. | Microsoft Defender for Endpoint – Possible coin mining activity – Trojan:Linux/CoinMiner!rfn Microsoft Defender for Cloud – Digital currency mining activity |
| Persistence | LiteLLM: SSH key written to a service account, cron entries created, and payload directories made immutable with chattr +i to resist cleanup. RAGFlow: api/__init__.py backdoored to load a hidden hook file on every service start, surviving container restarts. SSH key planted in the container. Kestra: miner launched with nohup to survive shell exit; follow-on harvest.sh collected and stored host data through the Kestra KV API. | Microsoft Defender for Endpoint – Suspicious addition of an SSH key; – Suspicious cron job creation; – Suspicious kernel module loaded |
| Command and Control | LiteLLM: outbound beacons to raw-IP infrastructure on port 81, sslip.io DNS rebinding to bypass reputation checks, and OAST callbacks to yosemite[.]jp, gobygo[.]net, and oast[.]me/pro/fun. RAGFlow: SSRF probing to shared scanning infrastructure in phase 1; API key exfiltration to a separate C2 endpoint in phase 2. Kestra: interactive reverse shell to a Linode VPS; sustained mining pool connections to auto.c3pool[.]org. | Microsoft Defender for Endpoint – Suspicious communication with a remote target; – Suspicious file or content ingress. – Suspicious connection to cryptocurrency mining pool |
Security Copilot customers can use the standalone experience to create their own prompts or run prebuilt promptbooks to automate incident response or investigation tasks related to this threat:
Microsoft Defender XDR customers can use these Advanced hunting queries to identify behaviors associated with this intrusion across Linux workloads and AI gateway environments. Each query focuses on a specific detection objective and is designed to help analysts validate suspicious activity, pivot across related process and network telemetry, and prioritize results that combine gateway-originated execution, secret access, payload staging, persistence, or outbound communication. Tune the queries for known administrative activity and approved gateway maintenance in your environment.
When reviewing results, prioritize events where a gateway process launches a shell, downloader, interpreter, or system utility; where command lines reference /proc/1/environ, LiteLLM database tables, provider keys, or PostgreSQL libraries; and where outbound traffic reaches raw-IP infrastructure or out-of-band callback domains. Matches that combine gateway ancestry, secret-access terms, and outbound communication should be treated as higher confidence.
This query looks for a LiteLLM gateway process launching execution utilities that are not expected for normal model-routing activity. In this intrusion, that relationship was the earliest high-value pivot: the gateway runtime became the parent process for shell commands, Python one-liners, downloaders, secret discovery, and follow-on payload execution.
// Low-FP pivot: AI gateway parent process spawning execution utilities.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine) and isnotempty(InitiatingProcessCommandLine)
| extend ParentCmd = tolower(InitiatingProcessCommandLine), Cmd = tolower(ProcessCommandLine)
| where ParentCmd has_any ("litellm", "litellm-proxy", "litellm_proxy", "ragflow", "kestra")
| where FileName in~ ("bash", "sh", "dash", "curl", "wget", "python", "python3")
| where Cmd has_any ("/proc/1/environ", "database_url", "psycopg2", "urllib.request", "urlretrieve", "base64")
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId
| sort by Timestamp asc
This query detects command-line access to /proc/1/environ, a high-signal behavior in containerized services where the main process often runs as PID 1. For an AI gateway, this environment can contain model-provider API keys, the gateway master key, database connection strings, UI passwords, and other secrets.
// High-signal secret access in containerized services.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| extend Cmd = tolower(ProcessCommandLine)
| where Cmd contains "/proc/1/environ"
| where FileName in~ ("cat", "bash", "sh", "python", "python3", "grep")
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId
| sort by Timestamp asc
This query narrows secret-discovery hunting to LiteLLM-specific context before matching sensitive terms. That structure reduces noise from generic words such as key, token, and password, while still surfacing command lines that reference LiteLLM proxy tables, virtual keys, provider configuration, or database material.
// Hunt for command lines that combine LiteLLM context with secret-related terms.
// This helps reduce false positives from generic credential keywords.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| extend Cmd = tolower(ProcessCommandLine)
| where Cmd has_any (
"litellm",
"litellm_proxymodeltable",
"litellm_verificationtoken",
"proxymodeltable",
"verificationtoken"
)
| where Cmd has_any (
"secret",
"token",
"key",
"password",
"master",
"database_url",
"postgres",
"psycopg2",
"psycopg2-binary"
)
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId, FolderPath
| sort by Timestamp asc
This query hunts for Python execution that references database connection material or PostgreSQL client libraries. In the observed attack chain, Python was used to parse DATABASE_URL, install or import PostgreSQL support, and access LiteLLM-backed database tables containing model configuration and virtual-key material.
// Hunt for Python activity associated with database credential discovery or use.
// Pivot from matches to parent process, network connections, and any package-install activity.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| extend Cmd = tolower(ProcessCommandLine)
| where Cmd has_any ("python", "python2", "python3")
| where Cmd has_any (
"database_url",
"postgres",
"postgresql",
"psycopg2",
"psycopg2-binary"
)
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId, FolderPath
| sort by Timestamp asc
This query looks for common Linux text-processing utilities used to search environment files, application configuration, or LiteLLM-related material for secrets. It requires three signals: a discovery utility, a relevant target, and a sensitive keyword, making it more precise than broad keyword searches alone.
// Hunt for shell utilities searching for secrets in environment or configuration data.
// Higher confidence results combine a discovery tool, a relevant target, and a secret keyword.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| extend Cmd = tolower(ProcessCommandLine)
| where Cmd has_any ("grep", "egrep", "fgrep", "awk", "sed", "cat", "strings")
| where Cmd has_any ("litellm", "database_url", "environ")
or Cmd contains "/proc/1/environ"
or Cmd contains ".env"
| where Cmd has_any ("secret", "token", "key", "password", "master", "postgres")
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId, FolderPath
| sort by Timestamp asc
This combined query is useful for triage dashboards or incident review because it labels each result with a detection reason. Analysts can use the DetectionReason field to quickly separate direct environment access, LiteLLM-specific secret discovery, Python database credential access, and shell-based searching.
// Combined triage query using only high-confidence secret-discovery signals.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| extend Cmd = tolower(ProcessCommandLine)
| extend DetectionReason = case(
Cmd contains "/proc/1/environ", "Direct access to PID 1 environment variables",
Cmd has_any ("litellm", "litellm_proxymodeltable", "litellm_verificationtoken", "proxymodeltable", "verificationtoken") and Cmd has_any ("database_url", "postgres", "psycopg2", "master", "secret", "token", "password"), "LiteLLM-related secret or database discovery",
FileName in~ ("python", "python2", "python3") and Cmd has_any ("database_url", "postgres", "postgresql", "psycopg2", "psycopg2-binary"), "Python-based database credential discovery",
"")
| where DetectionReason != ""
| project Timestamp, DeviceName, AccountName, FileName, DetectionReason, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId
| sort by Timestamp asc
This query identifies the payload-delivery pattern observed after gateway execution: raw-IP retrieval, staging under /tmp, and execution with supervisord-style arguments or bridge-related environment values. Review matches for masquerading, unexpected executable files in world-writable paths, and parentage from the gateway process.
// Hunt for staged payload execution and supervisord-style masquerading.
// Focus on /tmp execution, bridge variables, and known payload path fragments.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| where ProcessCommandLine has_any ("/private/python3", "/anonymus/bins_s", "BRIDGE_STANDALONE", "PORT")
or (FolderPath == "/tmp/python3" and ProcessCommandLine has "supervisord")
| project Timestamp, DeviceName, AccountName, FileName, FolderPath, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId
| sort by Timestamp asc
This query hunts for attempts to load the Linux msr kernel module with write access enabled. That behavior is strongly associated with performance tuning for RandomX/XMRig mining and is unusual on most production servers unless explicitly approved for low-level performance testing.
// Hunt for MSR write access often used to optimize RandomX/XMRig mining.
// Validate whether the host has any legitimate reason to load msr with allow_writes.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| where ProcessCommandLine has_all ("modprobe", "msr", "allow_writes")
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId
| sort by Timestamp asc
This query groups the persistence and defense-evasion behaviors observed in the intrusion: hidden-file relaunch from /tmp, cron manipulation, SSH authorized-key modification, and immutable-flag changes. These signals should be reviewed with process ancestry and file-write events to identify the account and payload responsible for durable access.
// Hunt for persistence and defense-evasion activity used to keep the payload running. // Review matches for service-account abuse, hidden /tmp execution, and cleanup resistance. DeviceProcessEvents | where isnotempty(ProcessCommandLine) | where (ProcessCommandLine contains "exec /tmp/." and ProcessCommandLine contains "-c /tmp/.") or (ProcessCommandLine contains "crontab" and ProcessCommandLine contains "grep -v") or ProcessCommandLine has "chattr" or (ProcessCommandLine has "authorized_keys" and ProcessCommandLine contains ">>") | project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId, FolderPath | sort by Timestamp asc
This query hunts for connections to infrastructure directly tied to the observed campaign. To reduce false positives, it focuses on known campaign domains/IPs and execution tools commonly used in the attack chain.
// Known campaign infrastructure only (low-FP network pivot).
DeviceNetworkEvents
| extend RU = tolower(RemoteUrl), RIP = tostring(RemoteIP)
| where RU has_any ("yosemite.jp", "gobygo.net", "auto.c3pool.org", "45.150.109.151.sslip.io")
or RIP in ("45.150.109.151", "135.125.10.56", "172.232.38.92", "47.86.197.116", "2001:41d0:701:1100::adfd")
| where InitiatingProcessFileName in~ ("bash", "sh", "dash", "python", "python3", "curl", "wget", "nohup")
| project Timestamp, DeviceName, InitiatingProcessAccountName, InitiatingProcessFileName, InitiatingProcessCommandLine, RemoteUrl, RemoteIP, RemotePort
| sort by Timestamp asc
For higher-confidence triage, correlate these results across time and telemetry types. A single match may represent administrative activity, but the combination of gateway-originated execution, secret access, database-focused Python, payload staging in /tmp, MSR tuning, persistence attempts, and outbound callbacks should be investigated as a potential end-to-end compromise path.
| Tactic | Technique | Observed activity |
| Initial Access | T1190 Exploit Public-Facing Application | Abuse of the internet-exposed LiteLLM gateway runtime |
| Execution | T1059 Command and Scripting Interpreter | python3 -c one-liners and shell scripts launched from the gateway process |
| Credential Access | T1552.001 Unsecured Credentials: Credentials in Files | Harvest of provider API keys from /proc/1/environ and the LiteLLM model-config table |
| Discovery | T1057 Process Discovery / T1518 Software Discovery | pgrep sweeps for rival miners and enumeration of PostgreSQL config files |
| Defense Evasion | T1036.005 Masquerading / T1564.001 Hidden Files and Directories | Payloads named after system daemons, executed from hidden /tmp files |
| Impact | T1496 Resource Hijacking | Cryptomining with MSR tuning and competing-miner eviction |
| Persistence | T1098.004 SSH Authorized Keys / T1053.003 Cron | Service-account SSH key and cron entries for durable access |
| Defense Evasion | T1222.002 Linux File and Directory Permissions Modification | chattr +i immutable flags on payload directories to resist cleanup |
| Command and Control | T1071.001 Application Layer Protocol / T1095 Non-Application Layer Protocol | HTTP beacons to raw-IP infrastructure, exfil to yosemite[.]jp, and OAST callbacks |
| IOC | Type | Role |
| 45.150.109[.]151 | IPv4 | Scanning/recon infrastructure – multiple targeted AI workloads |
| 135.125.10[.]56:19888 | IPv4:port | RAGFlow exploitation C2 — LLM API key exfiltration endpoint |
| 172.232.38[.]92:32991 | IPv4:port | Kestra reverse shell C2 (Linode VPS) |
| 45.150.109.151.sslip[.]io | Domain | DNS rebinding used in LiteLLM attacks to evade domain reputation checks |
| auto.c3pool[.]org:443 | Domain:port | XMRig Monero mining pool (Kestra) |
| 2001:41d0:701:1100::adfd | IPv6 | c3pool mining endpoint (Kestra) |
| 47.86.197[.]116 | IPv4 | c3pool mining endpoint (Kestra) |
| yosemite[.]jp | Domain | C2/exfiltration endpoint — LiteLLM credential harvesting (OAST + recv.php) |
| gobygo[.]net | Domain | C2 beacon infrastructure — subdomain-encoded LiteLLM beacons |
| oast[.]me / oast[.]pro / oast[.]fun | Domains | Out-of-band callback domains — execution confirmation and credential exfiltration (LiteLLM) |
| 194.213.18[.]133 | IPv4 | Attacker-controlled mail MX / mail infrastructure |
| File / Path | SHA256 | Notes |
| /tmp/d (ELF binary) | f64b88e9318bdf23f2dd119a0ce1dd1bdb3c8cd2e0e1e23ba3ef2e19072b79cc | LiteLLM #2 — unknown ELF; not on VirusTotal |
| XMRig cryptominer | 49fdcf32bfe837899a84e8938f0d07ae96ddd218a280a09eb60df8d64597bd8f | LiteLLM — XMRig binary |
| XMRig cryptominer | 3af9f25a4d45bb4f1ec5627cdbc6703cf3b4be75a892162d299d80ddfb266f42 | LiteLLM — XMRig binary (variant) |
| Installer / bridge script | 3d24ac736635e0fa0c5c459c9e18ca09d1ec9a1751a4503130934395609bd7e0 | LiteLLM — drops /tmp/python3 and launches supervisord bridge |
For the latest security research from the Microsoft Threat Intelligence community, check out the Microsoft Threat Intelligence Blog.
To get notified about new publications and to join discussions on social media, follow us on LinkedIn, X (formerly Twitter), and Bluesky.
To hear stories and insights from the Microsoft Threat Intelligence community about the ever-evolving threat landscape, listen to the Microsoft Threat Intelligence podcast.
Review our documentation to learn more about our real-time protection capabilities and see how to enable them within your organization.
The post When AI infrastructure becomes the target: Securing gateways and control points appeared first on Microsoft Security Blog.
The following article originally appeared on PulseMCP’s blog and is being republished here with the authors’ permission.
Most MCP demos feature a single server connecting to a single client. For example, you might wire up a Gmail MCP server to Claude Code. It works! It triages your inbox, drafts replies, finds that thing from three weeks ago.
But…it’s a little pointless. Gmail already has a perfectly good interface for email. It’s called Gmail. Google has spent 20 years refining it. If “chat with your inbox” were all MCP got you, you’d be right to wonder what the fuss is about.

Here’s the same Gmail MCP server doing something Gmail can’t do alone: filing business expense receipts.
An IKEA order confirmation is sitting in my inbox, receipt attached. Benepass—a benefits provider—wants that receipt uploaded and submitted. Claude Code can find the email, pull the attachment, log into Benepass by retrieving the one-time code from Gmail, upload the receipt, and submit. Done.
No single app could do this, because the job crosses apps. Server composition is a killer feature of MCP, not just “your AI talks to one integration.”

There is clear value in connecting multiple servers to a client like Claude Code. Taking the pattern a step further, there’s also value in connecting those same servers to other interfaces you like to use. Claude Code in your terminal shouldn’t be the only way you access AI. Claude.ai can be another entry point. Gmail can be another. Linear yet another. The list goes on.
You’ll often want the same integrations available when you’re doing a quick task inside Slack as you would while running a long agentic task with Claude Code.
You could be conversing with a coworker, CC in “@ai does our discussion here align well with current company strategy?” and get an inline response that incorporates your company OKRs from Notion and some recent Zoom transcripts from the last few product strategy meetings.
Or you’re working through your to-do list in Linear, realize you need to schedule a meeting as the action item for one of your tickets, and you can comment on a ticket “@ai schedule a meeting with John and draft an agenda based on the meeting I just had with Pam to address this.”
These native MCP client functionalities are officially on their way to many of the interfaces you use today. And we’re starting to see prominent nonnative integrations like Claude Tag fill this need too.
But even better: You can already set up yourself and your team with these capabilities today. You just need to bridge the desired surface to your own agent harness behind the scenes.

All it takes is awareness of the right patterns and some MCP-friendly glue.
Local servers are hard for end-users. “Just install npm, then edit this JSON file, then set these environment variables” isn’t something you want to say to most of your colleagues.
Remote servers are much easier: “Here’s a URL to paste in somewhere.” Some surfaces only support remote servers, like claude.ai.

So you can use local servers to prove out some flow, but you should prefer remote. Many servers you may want to use are only provided as local implementations, so you need a tool to convert them to remote servers. Because most local servers were designed to be used by one authenticated user at a time, this can be tricky.
The pattern that solves this pain point is to use a bridge like mcp-auth-wrapper to take a local server and turn it into a remote, multitenant one: OAuth in front, per-user credentials behind. Sharing any server becomes sharing a link.
Okay, we’re down to one link per server. But exactly where do my users put this link?
The one-size-fits-all pattern here is to provide installation guidance for every possible MCP client app. A service like install-mcp makes this easy: clean UX to generate the right install link and per-client instructions, so you share one page instead of writing a tutorial.
Or if you want to roll it into your own interface, tap into the TypeScript library version of this: mcp-install-instructions.
We’re working on standardizing the mcp.json file format here to simplify this story in the long term, but ultimately we still expect non-CLI interfaces to each have a slightly different flow for “enabling a connector.” For many enterprises, this responsibility may soon become solely an IT affair with enterprise-managed auth.
What happens when we start to bring new clients into the fold? By default, each user has to configure and authenticate each MCP server independently. 10+ “add connector” flows. 10+ OAuth dances. Done once for Claude Code, then again for Linear, for email…and so on.
The solution: Centralize your configuration and auth storage in a single, aggregated MCP server. A deployment of mcp-aggregator can do this for you: one endpoint, one login, every tool namespaced behind it. Configure each client once. Login once per service. Share one link with your teammates.

mcp-aggregator is the minimal, DIY version of an MCP gateway. If you have enterprise needs like SSO integration or fine-grained IT admin controls, you could use this layer of MCP server consolidation as the piece of your infrastructure where you can add those bells and whistles.
By default, mcp-aggregator also loses out on the ability to use MCP server configurations as per-session constraints. However, you could build out your mcp-aggregator deployment to re-enable the constraints that matter to you. For example, by offering query parameters like ?readOnlyTools=true or ?serversEnabled=datadog,sentry. Or by giving discrete MCP servers unique endpoints (but still aggregating the auth concerns under the hood).
Most of your MCP server needs can be solved by walking through the options for choosing (or building) an MCP server. But is it always worth the effort?
Say you switched home utilities plans. An energy provider, Octopus Energy, emailed you a confirmation. You want the pricing data in Home Assistant, which controls your smart thermostat, and Octopus Energy does have an API for it—but the API key is behind a website login, and there’s no API for getting the API key.
Solution: Give the agent a screen. computer-use-mcp is a minimal example of an MCP server that can click and type like a human to log into the website, copy the key, then go back to clean API calls to set up the dashboard.
With that one escape hatch, “there’s no API for that” or “it’s too much work to find an MCP server” stops being a blocker forever.
To borrow from Anthropic’s framing of this problem:
If you have an MCP server for the service, Claude uses that.
If the task is a shell command, Claude uses Bash.
If the task is browser work and you have Claude in Chrome set up, Claude uses that.
If none of those apply, Claude uses computer use.
And indeed, if you find yourself regularly leaning on inefficient approaches like browser use or computer use to get a predictable workflow done, it may be worth your while to abstract away those Playwright calls into a reliable MCP server, like this Good Eggs example.
Some servers only work when they run on your own personal machine. computer-use is the obvious example. And sometimes, you just want to use your phone to tell Claude to do something that relies on something readily available on your laptop upstairs.
mcp-local-tunnel shows off the pattern that makes machine-bound servers available through remote aggregation just like everything else.

Naive tool calling has a much-maligned drawback: Every intermediate result flows through the model. And for many clients, every tool definition is placed in context up front, meaning you’ve spent tens of thousands of tokens before even starting.
For example, asking the agent to copy a doc’s contents from Google Drive into Salesforce means the entire doc passes through the model twice—for no reason. Anthropic measured one workflow dropping from 150,000 tokens to 2,000 by fixing this.
There are several ways to solve this problem, each with its own trade-offs worth its own blog post:
tool-sandbox-mcp shows how a single execute_code abstracts away the problem by adding a layer of code execution in between your agent and your aggregated set of servers. Works inside any MCP client.call-mcp shows how wrapping your tools this way means you can compose them with your usual CLI toolkit—shell, jq, cron, scripts, and other CLI tools.
With your MCP aggregator in hand, you can now connect any service to any other service with just a little bit of glue code. When each service starts to implement an MCP client natively, this will be done for you. But for now, you can tackle the opportunities application by application.
We’ll use a Linear integration as our example, and you can imagine doing almost the exact same thing with any project management tracker like Jira, Asana, and so on.
The goal: inject a highly-capable MCP client, powered by your favorite coding agent, into a workflow inside a SaaS application you already use. We’ll take advantage of the fact that we have already consolidated all our integrations in the mcp-aggregator pattern above.
Here’s an end-to-end project showing off how to build this sort of “Linear harness” that bolts a Claude Code setup to serve as the MCP client.
You can see how:

Going beyond Linear and task management, here’s a Claude Code-powered agent that lives inside Minecraft in a few hundred lines. In game, we ask it to:
It’s silly on purpose, but the point is that a fully capable agent—all our tools, all these patterns—can run anywhere with a little glue. And it’s all attainable today, for individuals, for teams, and for enterprises alike.