“100% of my code is written by [insert whichever AI coding agent is popular right now]!” You’ve probably heard this claim many times this year.
We wanted to find out how many developers have actually fully outsourced code writing to agents and how the share of agent-generated code differs across regions, tech stacks, and seniority levels.
Luckily, our large-scale, globally representative Developer Ecosystem Survey 2026 gave us the perfect opportunity to uncover the trends. In May–July 2026, we asked over 15,000 professional developers worldwide: “What percentage of the code that you produced last month for work was …?
Fully generated by AI agents
Written by you with some AI assistance
Fully written by you without any AI assistance”
The answer options were: 0%, 1%–20%, 21%–40%, …, 81%–99%, 100%, and I don’t know.
Here is what we found:
Expectedly, manual coding is disappearing quickly, but not everybody has made the leap to a fully agentic development workflow yet.
On average, professional developers report that:
~47% of their code is fully written by agents.
~38% is written with some AI assistance.
~27% is written fully manually.
However, adoption varies wildly across the board. Over half of all developers now write less than 20% of their code manually, and one in five writes literally zero code without AI help. At the same time, the group that relies almost entirely on coding agents (over 80% agent-generated code) remains a minority of around 22%. Most developers are sitting somewhere in the middle.
Agentic coding by professional experience
Interestingly, senior developers are among the first to hand coding over to agents. About a quarter of senior developers generate the vast majority of their code (over 80%) using agents, compared to a smaller fraction of juniors who tend to lean more toward AI-assisted workflows rather than fully agentic coding.
That said, not all seniors are agentic-first yet. Adoption varies widely within this group. See the charts below.
Agentic coding by most-used AI coding tools
About 32% of developers who report Claude Code as their most-used AI coding tool generate over 80% of their code with agents.
Interestingly, among developers who use Codex, this share is notably higher at 42%. The share of developers who don’t write code without AI assistance at all is 37% among Codex users, which is tangibly higher than among users of other AI coding tools.
Our interpretation is that while Claude Code is increasingly becoming the mainstream AI coding tool (already the de facto standard, with 39% adoption at work), its audience no longer consists predominantly of advanced users of agents. Codex, on the other hand, is catching up in terms of awareness and adoption, and its user base might have more advanced users seeking better value for money (Codex has historically offered higher quotas).
Cursor users are similar to Claude Code users – on average, 58% of their code is agent-generated, and for 28% of its users, coding agents generate over 80% of their code.
Agentic coding by main programming language
There is a clear split in agentic coding adoption by tech stack. Developers with Go, JavaScript, and TypeScript as their main programming languages report the highest shares of agent-generated code, averaging 54%–55%.
On the other end of the spectrum, C and C++ developers remain the least agentic, maintaining a much higher proportion of manually written code – 38% on average.
Java and Python developers sit in the middle with 48%–51%.
Agentic coding by regions
The geographic differences are remarkable. Developers across East Asia, particularly in China, Japan, and South Korea, are leading the charge in agentic coding: About twice as many developers there (32%–35%) generate the vast majority of their code (over 80%) with agents, compared with around 16% of developers in Europe and the UK.
Segments of developers by AI usage
We were curious whether developers could be grouped into homogeneous segments based on how they write code today. Three distinct profiles emerged:
See the methodology notes for more details on how we did this.
We deliberately use the word “coders” here to highlight that this is about the code generation process, not development as a whole.
Agentic coders (~31% of developers) mostly write code with agents nowadays:
On average, 84% of their code is fully agent-generated.
~15% is written with some AI assistance.
~6% is written manually without AI at all.
Despite being the locomotives of agentic coding, only 46%–57% of heavy users of Claude Code and Codex are agentic coders.
AI-assisted coders (~47% of developers) haven’t gone full agentic yet, but already write a large portion of their code with agents. They still prefer an AI-assisted development workflow and don’t hesitate to write some code manually if needed:
On average, 40% of their code is fully agent-generated.
~60% is written with AI assistance.
~20% is written manually.
Manual coders (~23% of developers) write most of their code manually, though they use AI sometimes, mostly in an AI-assisted rather than fully agentic manner:
On average, ~10% of their code is agent-generated.
~25% is AI-assisted.
~75% is manually written.
Whether you are all-in on agentic workflows or prefer hands-on coding, the shift is clear: Purely manual coding is quickly becoming a thing of the past.
In our previous blog post, based on the same survey data, we explored trends in AI coding agent adoption across the industry and the main players in this market.
We plan to share more materials on agentic development from the Developer Ecosystem Survey 2026 with the community soon.
Stay tuned and subscribe to JetBrains Research blog updates below!
Methodology notes
Unrealistic responses – where the sum of the lower bounds of selected answers exceeded 150% or the sum of the upper bounds of selected answers fell below 80% – were not used in this analysis. This filter was applied on top of the regular data-cleaning filters used for Developer Ecosystem Survey data.
We used the midpoint of each answer bucket (0%, 10.5%, 30.5%, 50.5%, 70.5%, 90%, 100%) to calculate averages. The averages across the three categories of how code is written within the same group (e.g. seniors) could exceed 100% because of the bucketed nature of the answers, and respondents’ self-reports may not always be fully accurate.
We employed hierarchical cluster analysis with the Ward method and Euclidean distance on unstandardized bucket midpoints (e.g. 0%, 10.5%, 30.5%) to clusterize developers into homogeneous segments based on how they write code.
In this report, “professional developers” refers to respondents who reported being involved in coding or programming in any of the following job roles:
Developer / Programmer / Software Engineer
AI / ML Engineer
DevOps Engineer / Infrastructure Developer
Architect
Data Scientist / Data Engineer / Data Analyst
QA Engineer
Roughly 90% of the sample falls into the Developer / Programmer / Software Engineer job category.
The Developer Ecosystem Survey is localized into eight languages: English, Spanish, Chinese, Japanese, Korean, German, French, and Portuguese. We apply quotas on the required number of responses by region to help achieve accurate global representation. The quotas are proportionate to the number of developers in each region, based on estimates by our Data Science team. The detailed methodology of these estimates is described here.
The Developer Ecosystem Survey has been statistically reweighted to better represent the global developer population by region, employment status, programming language, and familiarity with JetBrains products (to avoid data skewed toward an excessively JetBrains-familiar audience). You can read about the weighting methodology for the Developer Ecosystem Survey here.
AI is creating a new layer of enterprise infrastructure. Gateways, retrieval platforms, orchestration services, and containerized runtimes now sit between users, applications, data, and models. These systems concentrate credentials, data access, model connectivity, and execution privileges, making them some of the most powerful components in the AI stack.
That concentration of trust is also creating new opportunities for attackers. In recent investigations, Microsoft observed activity targeting three distinct AI workloads: a LiteLLM gateway, a RAGFlow deployment, and a Kestra workflow environment. The intrusion paths varied, but the objectives were strikingly similar. Attackers sought to steal credentials, establish persistence, and monetize compromised compute resources.
The individual techniques matter, but the broader pattern matters more. Across these cases, attackers treated AI infrastructure as a control plane where credential theft, host compromise, and downstream data access can converge. As organizations continue to deploy AI systems, these platforms are becoming high value targets that deserve the same security scrutiny as other critical enterprise infrastructure.
AI workloads are becoming high-value control points
The campaign-level signal extends beyond one product. The targeted workloads served different functions, but each exposed assets that could support follow-on abuse, including model-provider keys, proxy-issued virtual keys, database connection strings, tenant configuration, workflow execution, or host compute. Post-compromise behavior varied by workload role. Defenders should inventory exposed AI management surfaces, restrict administrative access, and monitor for gateway-originated execution and secret access.
Three observed compromises across AI workloads
AI workload
Observed activity
Attacker objective
LiteLLM
Observed attacker activity: Python droppers, runtime secret harvesting, PostgreSQL collection, miner deployment, and persistence activity from the LiteLLM gateway context.
Microsoft assessment: Initial access likely occurred through exploitation of the exposed LiteLLM gateway surface, consistent with the vulnerability chain involving CVE-2026-42271 and CVE-2026-48710.
Observed attacker activity: Possible SSRF-style reconnaissance followed several days later by code execution, application-path modification, and placement of a Python hook in the TenantLLM credential-configuration flow.
Public research: Describes multiple RAGFlow execution paths; Microsoft does not attribute this intrusion to a specific vulnerability.
Intercept newly configured LLM provider credentials and model metadata.
Kestra
Observed attacker activity: Workflow-origin shell execution, Docker and container-environment discovery, XMRig deployment, and follow-on data collection.
Microsoft assessment: Initial access likely involved exploitation of the exposed Kestra orchestration surface, with CVE-2026-49869 providing relevant public vulnerability context.
Secret discovery, container-level access, data collection, and rapid compute monetization.
Case study 1: LiteLLM gateway compromise
Framework role and affected runtime context
LiteLLM is commonly deployed as a proxy or gateway between applications and model providers. In that position, the service may hold or retrieve model-provider keys, LiteLLM master keys, virtual-key records, database connection strings, routing configuration, and tenant policy data. Command execution in the gateway runtime therefore exposed a process context close to AI routing and credential material.
Microsoft assesses with high confidence that initial access likely occurred through exploitation of the exposed LiteLLM gateway surface. Relevant public vulnerability paths include CVE-2026-42271, an authenticated command-execution issue in LiteLLM MCP stdio test endpoints, and the route described in public research that chains this flaw with CVE-2026-48710, a Starlette host-header validation bypass, to achieve unauthenticated remote code execution in vulnerable exposed deployments.
In this chain, CVE-2026-42271 provides the command execution capability through the MCP stdio test path, while CVE-2026-48710 can weaken the authentication boundary in affected configurations, potentially making that capability reachable without valid credentials.
In this case, initial access occurred in the context of the LiteLLM gateway process. The gateway service, rather than an unrelated system process, became the execution origin. Subsequent activity from that point is described in the observed attack chain below.
Figure 2. Process tree observed from the compromised LiteLLM gateway, showing shell and Python execution originating from the gateway service process.
Observed attack chain
Stage 1: Credential harvesting from the gateway runtime
The first observed stage was credential harvesting from the LiteLLM gateway runtime. The payload read the gateway process environment and filtered for credential-related values, including model-provider API keys, the LiteLLM master key, database connection strings, UI credentials, tokens, passwords, and other secret-like fields.
Figure 3. Credential harvesting from the gateway process environment, filtered for provider keys and connection strings.
In containerized LiteLLM deployments where the gateway runs as PID 1, /proc/1/environ exposes the environment block for the gateway process. Telemetry showed the payload reading /proc/1/environ, filtering for keywords such as master, API key, token, password, and UI-related fields, then sending collected values to attacker-controlled infrastructure.
The exfiltration logic used multiple transports in sequence, including Python urllib, curl, and wget. This provided fallback paths if one tool was unavailable or if egress controls affected one outbound method.
Stage 2: Payload delivery and masqueraded execution
The second stage moved from gateway-level command execution to payload delivery. The first delivery path launched from the compromised LiteLLM gateway process as an inline Python command. The code retrieved a masqueraded ELF binary from attacker-controlled infrastructure, staged it under a temporary path, marked it executable, and launched it with command-line arguments resembling a Linux service process.
The downloaded ELF used service-style naming and arguments to masquerade as a benign Linux daemon.
A second delivery path used a shell-stage downloader. A gateway-spawned Python command invoked a shell that used multiple download methods with short timeouts and fallback behavior, staged the retrieved content under randomized temporary paths, marked it executable, and launched it with supplied parameters. Together, these paths show redundant payload retrieval and execution from the gateway process context.
Figure 4. ELF binary retrieved and staged under the interpreter’s name python3, then launched with service-manager argumentsStage 3: Host discovery and competing-miner checks.
The third stage performed host discovery from the second-stage payload. Observed commands fingerprinted the host, checked privilege boundaries, inspected listening ports, and searched for other miner or remote-access activity on the system.
Figure 5. Host reconnaissance and competing-miner sweeps.
Relevant artifacts included a silent passwordless-sudo check, listening-port inspection, and multiple process sweeps for miner or remote-access tooling. The searches included terms associated with miners, remote shells, competing infrastructure, and previously dropped scripts.
Stage 4: Cryptomining preparation and competing-miner removal
The fourth stage prepared the host for cryptocurrency mining and removed competing persistence. Telemetry showed XMRig or XMRig-like mining components, RandomX-related tuning, and cleanup commands targeting other miners.
Figure 6. MSR module loaded for CPU tuning, followed by removal of competing miner cron entries.
Two artifacts were central. A command loaded the Linux Model-Specific Register module with write access enabled, a behavior commonly associated with RandomX/XMRig CPU tuning. Separately, a crontab rewrite removed entries associated with other miner names before writing the modified cron data back.
Stage 5: LiteLLM database access through Azure PostgreSQL
The fifth stage used the previously collected database connection string to access the LiteLLM-backed PostgreSQL data tier. This moved the activity from host-level secret exposure into application-layer data collection from the gateway backend database.
Figure 7. Discovery of PostgreSQL configuration files and native-extension paths.Figure 8. Database access and credential collection from LiteLLM model and virtual-key tables.
Telemetry captured a self-contained python3 one-liner that read the gateway PID 1 environment, parsed DATABASE_URL, installed or imported PostgreSQL support, connected to the database, and dumped LiteLLM model and virtual-key records. The connection string referenced <name>.postgres.database.azure.com, indicating Azure Database for PostgreSQL.
The payload targeted LiteLLM tables including LiteLLM_ProxyModelTable and LiteLLM_VerificationToken. The collected records could include model configuration, upstream provider key material, provider endpoints, and proxy-issued virtual keys.
Output was base64-encoded and exfiltrated in small chunks to an out-of-band callback endpoint. A sibling variant posted data to a separate web endpoint that was also observed during the earlier credential-harvesting stage.
Stage 6: Persistence, command-and-control, and defence evasion
The sixth stage added persistence, command-and-control, and defence-evasion mechanisms. Observed artifacts included service-account SSH authorized-key modification, hidden-file relay execution, masqueraded service names, self-relaunch loops, and immutable-file attributes.
The durable access artifact was an authorized_keys write under a service account. Additional artifacts included hidden-file relay execution, command-and-control relay components, masqueraded systemd service names, and relaunch paths under hidden temporary files.
Names used in relaunch paths overlapped with common Linux daemon naming patterns. Periodic out-of-band callbacks were also observed, providing network telemetry that the payload continued to execute and retained outbound connectivity.
Impact
The LiteLLM compromise produced multiple impact paths: provider credential exposure, proxy-issued key exposure, database-backed configuration access, host resource abuse, and durable service-account access. The gateway role made these impacts broader than a standard single-process application compromise.
Case study 2: RAGFlow compromise
Framework role and affected runtime context
RAGFlow supports document-processing and retrieval-augmented generation workflows and stores tenant LLM configuration. The observed execution occurred inside the RAGFlow container under the application runtime lineage. That context is important because the affected code paths process provider credentials when users add or modify LLM settings.
Initial access and compromise pattern
Figure 10. RAGflow compromise – attack chain.
Microsoft assesses with high confidence that initial access likely occurred through exploitation of the exposed RAGFlow application surface. Telemetry showed the RAGFlow server process retrieving an attacker-supplied URL through the application’s own HTTP client, resulting in an outbound Burp Collaborator callback without corresponding child-process execution. Remote code execution in the same service context followed later in the observed sequence.
Microsoft assesses with low confidence which specific vulnerability, if any, enabled that code execution. Because the relevant application code paths execute within the RAGFlow Flask service process, endpoint telemetry could not distinguish the precise execution sink. Publicly documented vulnerabilities affecting relevant RAGFlow versions include CVE-2026-45312 and CVE-2026-28797, authenticated Jinja2 server-side template injection issues in the prompt generator and Agent workflow components; CVE-2026-24770, a MinerU parser path-traversal issue that can permit arbitrary file overwrite and subsequent code execution; and CVE-2025-68700, a Canvas CodeExec sandbox-bypass issue tracked as GHSA-8xw3-v6c2-j84j.
These vulnerabilities provide plausible technical context but are not attributed as the confirmed cause of this intrusion. Depending on the affected version and deployment configuration, access to authenticated functionality could also be influenced by separate account-access weaknesses, including CVE-2025-69286. For defenders, the possible SSRF activity through the OASTify relay network is a useful precursor signal because remote code execution in the same service context followed several days later.
Observed attack chain
Stage 1: Application discovery and hook creation
The first payload stage located the RAGFlow installation from inside the container and identified the tenant LLM model-configuration path. Telemetry showed discovery logic for common application locations, followed by creation of a hidden runtime hook under the application tree.
Figure 11. First stage Python credential theft hook.
Stage 2: Persistence through application startup modification
The second stage modified the application startup or import path so the hidden hook would load with the RAGFlow service. This tied the credential-interception behavior to the application runtime rather than to a separate long-running process.
Figure 12. Exec hook created in the startup path of RAGFLow.
Stage 3: Credential interception during LLM configuration
The hook wrapped the tenant LLM configuration flow and captured newly supplied provider metadata during credential setup. Captured fields included provider type, model name, API key material, and related endpoint metadata. The collection routine used outbound HTTP from within the container and suppressed errors so the application flow could continue if collection failed.
Figure 13. Credential Stealer extracting configured API keys.
Stage 4: Finalization and installation verification
The final stage wrote or refreshed the hook and created a local marker indicating that installation had completed. Command-line telemetry was partially truncated, but the repeated execution sequence, process lineage, and application-file modifications were sufficient to reconstruct the functional behavior.
Figure 14. Exfiltration of collected data to C2.
Impact
The RAGFlow compromise was primarily focused on LLM credential collection rather than host monetization. Telemetry did not show miner deployment or an interactive reverse shell in this case. The affected runtime path could capture provider credentials configured after the hook was installed, and the startup-path modification could persist across service restarts if the modified filesystem state remained present. SSH-key material was also written inside the container, but its durability depends on container privileges, filesystem persistence, and host-container boundary configuration.
Case study 3: Kestra compromise
Framework role and affected runtime context
Kestra is a workflow orchestration environment. Because workflows are designed to execute tasks and interact with external systems, abuse of workflow-creation and execution capabilities can provide direct code execution in the worker runtime.
Initial access and compromise pattern
Figure 15. Kestra Compromise – attack chain.
Microsoft assesses with high confidence that initial access likely occurred through exploitation of CVE-2026-49869, a critical authentication-bypass vulnerability in Kestra. Exploitation could allow an unauthenticated remote attacker with network access to bypass the login mechanism, define a malicious workflow using the Process runner, and trigger worker-side shell-script execution.
Following the assessed initial-access sequence, telemetry showed two closely timed workflow-origin shell sessions. The first produced shell initialization activity, while the second performed the main follow-on actions, including Docker socket access, container-environment enumeration, miner deployment, and defence-evasion file operations. A later workflow-origin event used a curl-pipe-shell delivery pattern to retrieve remote script content directly into a shell and store collected output through the application’s own key-value interface.
Observed attack chain
Stage 1: Workflow-origin shell execution
Telemetry showed the Kestra worker lineage spawning shell activity from the orchestration layer. Two closely timed workflow-origin shell sessions were observed; the first produced shell initialization activity, while the second performed the main follow-on actions.
Stage 2: Docker container environment discovery
After workflow-origin execution, commands accessed the mounted Docker socket from inside the compromised orchestration environment. The activity queried container metadata and inspected container environment arrays, exposing environment-backed values from other containers reachable through the mounted runtime socket.
This behavior is significant because workflow engines often run near automation secrets. Environment arrays, mounted configuration, service credentials, and container metadata may expose cloud keys, database passwords, API tokens, or internal service endpoints when the container runtime socket is accessible.
Figure 16. Container discovery performed through malicious workflow.
Stage 3: Cryptominer deployment
The monetization phase followed the workflow-origin execution chain. Telemetry showed miner retrieval from a public release source, archive extraction, binary renaming, background execution, and mining-pool communication. CPU-tuning behavior commonly associated with RandomX/XMRig mining was also observed.
Additional defence-evasion file operations were observed around a temporary path, including restrictive permissions and immutable-file attributes. These artifacts provide file-system telemetry alongside the workflow-origin process lineage and network activity.
Figure 17. Credential harvesting performed through malicious workflow.
Stage 4: Data harvesting through workflow task execution
A later workflow-origin event used a curl-pipe-shell pattern for follow-on collection. Remote script content was retrieved and executed directly by the shell without being written as a standalone script file first. The resulting output was encoded and stored through Kestra’s own key-value interface.
Figure 18. Deployment of cryptominer through malicious workflow.
Impact
The Kestra compromise exposed four impact paths: shell execution through the workflow engine, container-environment exposure through Docker socket access, host resource hijacking through miner deployment, and follow-on collection through workflow task execution. The later curl-pipe-shell event encoded collected output and stored it through Kestra’s own key-value interface, reducing reliance on standalone file artifacts.
Possible AI-assisted payload development
Several payloads exhibited characteristics often associated with assisted or generated code, including organized imports, explicit timeout handling, dependency fallbacks, formatted output, defensive exception handling, and explanatory comments. Compared with minimal, one-off shell payloads, these samples showed a more structured and robust implementation style.
Figure 19. Dropper source with structured imports, timeout handling, and non-English comments.Figure 20. Collection routine with dependency fallback on import failure.
These characteristics are observations about the tooling, not evidence of attribution. From a security perspective, their significance is that they can improve payload portability and resilience across Linux and container environments. No conclusion about the code’s authorship or development method is required.
Key patterns observed across AI workloads
Initial access differed by workload. LiteLLM involved command execution from the gateway runtime. RAGFlow progressed from SSRF-style probing to runtime modification. Kestra used workflow execution as the shell-access path.
The observed objectives were consistent. Across the cases, telemetry showed credential collection, durable access mechanisms, and resource monetization, even though the execution path differed by product.
Payload behavior was specific to each workload. LiteLLM payloads targeted gateway environment variables and database-backed proxy records. RAGFlow activity targeted LLM credential configuration. Kestra activity focused on workflow execution, container discovery, and cryptomining.
What this means for defenders: Defenders should monitor AI workloads according to their control-plane role, not only as isolated applications. Gateway, retrieval, and orchestration services can concentrate credentials, database access, workflow execution, and container privileges in one runtime. High-value detections should therefore correlate unexpected application-origin shells or interpreters with secret access, application-file modification, Docker socket use, outbound callbacks, and resource-hijacking activity. Treating these signals as a connected compromise path can expose attacks earlier than product-specific indicators alone.
Mitigation and protection guidance
Microsoft recommends the following mitigations to help reduce the risk and impact of AI workload compromise.
Treat AI gateways as Tier-0 secrets stores. Keep LiteLLM and similar proxies patched, require authentication across API and UI surfaces, restrict administrative and management ports, and do not expose management interfaces directly to the internet.
Scope and protect provider credentials. Issue per-team virtual keys with spend limits instead of sharing master keys, store upstream API keys in a managed secret store rather than process environment variables, and rotate credentials associated with an exposed or compromised gateway.
Apply least privilege to gateway and database access. Run the proxy under a dedicated service account, limit its PostgreSQL permissions to required objects, place the database behind a private endpoint with restrictive firewall rules, and enable Microsoft Defender for Cloud monitoring for the database and surrounding cloud resources.
Constrain outbound traffic. Use deny-by-default egress rules and allowlist only required model-provider and service endpoints. Block direct connections to raw-IP hosts and non-standard ports, and route permitted traffic through an FQDN-filtering firewall or inspecting proxy.
Monitor outbound callbacks and campaign infrastructure. Filter and log DNS traffic to identify out-of-band callbacks and subdomain-encoded beacons, and monitor connections to campaign-associated C2 and OAST domains.
Harden the host runtime. Mount temporary directories as non-executable where operationally feasible, alert on execution from world-writable paths, and monitor changes to cron entries, SSH authorized_keys files, and immutable-file attributes.
Enable Microsoft Defender for Endpoint protections on Linux. Keep real-time and cloud-delivered protection enabled to detect files written to disk, newly observed droppers, miners, and second-stage payloads. Enable behavior monitoring for anomalous child processes, credential access, data staging, exfiltration, and persistence activity.
Microsoft Defender detections
Microsoft Defender coordinates detection, prevention, investigation, and response across endpoints, identities, cloud workloads, and apps to provide integrated protection against attacks on AI infrastructure like the one discussed in this blog. Given the criticality of this new attack layer, defender is providing differentiated visibility, detection and protection from attacks against AI resources. Customers with provisioned access can also use Microsoft Security Copilot in Microsoft Defender to investigate and respond to incidents, hunt for threats, and protect their organization with relevant threat intelligence.
Tactic
Observed activity
Microsoft Defender coverage
Initial Access
Exploitation of internet-exposed AI workload surfaces, including model gateway, retrieval, and workflow orchestration services reachable without network restriction.
Microsoft Defender for Endpoint – Suspicious shell execution from an AI workload process – Suspicious shell execution from a scripting application runtime
Credential Access
LiteLLM: reads of /proc/1/environ and the model-config table to harvest provider API keys and the database connection string.
RAGFlow: TenantLLM.insert() monkey-patched to intercept provider API keys (OpenAI, Azure, Anthropic, Gemini) on every LLM configuration event, exfiltrated to a secondary C2 endpoint.
Kestra: Docker socket used to enumerate container Config.Env arrays across all running containers, collecting embedded cloud, database, and API secrets.
Microsoft Defender for Endpoint – Suspicious process collected data from local system – Suspicious file copy operations Enumeration of files with sensitive data
Execution & Defense Evasion
LiteLLM: second-stage binary dropped to /tmp and executed under names impersonating system services and daemons.
RAGFlow: base64-encoded Python payloads decoded and written to /tmp, executed sequentially to discover the RAGFlow install, inject a persistence hook, and verify implant success — fully automated with no interactive shell.
Kestra: malicious workflow submitted via the pipeline API caused the Java worker to spawn a bash reverse shell; XMRig was downloaded, unpacked, and renamed to evade name-based detection.
Microsoft Defender for Endpoint – Hidden file executed – Suspicious process launched from a world-writable directory – Suspicious path deletion – Suspicious file dropped and launched – Suspicious shell command execution – Suspicious piped command launched – Executable permission added to file or directory Possible reverse shell – Suspicious Python command-line execution\ – Suspicious script launched – Process launched in the background – Suspicious file or information obfuscation detected – Suspicious deletion of launched process binary – Suspicious shell execution from a scripting application runtime
RAGFlow: every LLM API key configured after infection silently exfiltrated, enabling unauthorized use of provider accounts at the attacker’s direction.
Kestra: XMRig v6.26.0 launched with RandomX MSR tuning toward a Monero mining pool, consuming host CPU for attacker profit.
Microsoft Defender for Endpoint – Possible coin mining activity – Trojan:Linux/CoinMiner!rfn
Microsoft Defender for Cloud – Digital currency mining activity
Persistence
LiteLLM: SSH key written to a service account, cron entries created, and payload directories made immutable with chattr +i to resist cleanup.
RAGFlow: api/__init__.py backdoored to load a hidden hook file on every service start, surviving container restarts. SSH key planted in the container.
Kestra: miner launched with nohup to survive shell exit; follow-on harvest.sh collected and stored host data through the Kestra KV API.
Microsoft Defender for Endpoint – Suspicious addition of an SSH key; – Suspicious cron job creation; – Suspicious kernel module loaded
Command and Control
LiteLLM: outbound beacons to raw-IP infrastructure on port 81, sslip.io DNS rebinding to bypass reputation checks, and OAST callbacks to yosemite[.]jp, gobygo[.]net, and oast[.]me/pro/fun.
RAGFlow: SSRF probing to shared scanning infrastructure in phase 1; API key exfiltration to a separate C2 endpoint in phase 2. Kestra: interactive reverse shell to a Linode VPS; sustained mining pool connections to auto.c3pool[.]org.
Microsoft Defender for Endpoint – Suspicious communication with a remote target; – Suspicious file or content ingress. – Suspicious connection to cryptocurrency mining pool
Microsoft Security Copilot
Security Copilot customers can use the standalone experience to create their own prompts or run prebuilt promptbooks to automate incident response or investigation tasks related to this threat:
Incident investigation: correlate gateway process, credential-access, mining, and persistence signals into a single timeline and surface the provider keys that may have been exposed.
Microsoft user analysis: assess accounts and service principals whose credentials the gateway could have exposed.
Advanced hunting queries
Microsoft Defender XDR customers can use these Advanced hunting queries to identify behaviors associated with this intrusion across Linux workloads and AI gateway environments. Each query focuses on a specific detection objective and is designed to help analysts validate suspicious activity, pivot across related process and network telemetry, and prioritize results that combine gateway-originated execution, secret access, payload staging, persistence, or outbound communication. Tune the queries for known administrative activity and approved gateway maintenance in your environment.
When reviewing results, prioritize events where a gateway process launches a shell, downloader, interpreter, or system utility; where command lines reference /proc/1/environ, LiteLLM database tables, provider keys, or PostgreSQL libraries; and where outbound traffic reaches raw-IP infrastructure or out-of-band callback domains. Matches that combine gateway ancestry, secret-access terms, and outbound communication should be treated as higher confidence.
AI gateway process spawning shells, downloaders, or interpreters
This query looks for a LiteLLM gateway process launching execution utilities that are not expected for normal model-routing activity. In this intrusion, that relationship was the earliest high-value pivot: the gateway runtime became the parent process for shell commands, Python one-liners, downloaders, secret discovery, and follow-on payload execution.
// Low-FP pivot: AI gateway parent process spawning execution utilities.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine) and isnotempty(InitiatingProcessCommandLine)
| extend ParentCmd = tolower(InitiatingProcessCommandLine), Cmd = tolower(ProcessCommandLine)
| where ParentCmd has_any ("litellm", "litellm-proxy", "litellm_proxy", "ragflow", "kestra")
| where FileName in~ ("bash", "sh", "dash", "curl", "wget", "python", "python3")
| where Cmd has_any ("/proc/1/environ", "database_url", "psycopg2", "urllib.request", "urlretrieve", "base64")
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId
| sort by Timestamp asc
Direct access to container environment variables
This query detects command-line access to /proc/1/environ, a high-signal behavior in containerized services where the main process often runs as PID 1. For an AI gateway, this environment can contain model-provider API keys, the gateway master key, database connection strings, UI passwords, and other secrets.
// High-signal secret access in containerized services.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| extend Cmd = tolower(ProcessCommandLine)
| where Cmd contains "/proc/1/environ"
| where FileName in~ ("cat", "bash", "sh", "python", "python3", "grep")
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId
| sort by Timestamp asc
LiteLLM-specific secret and configuration discovery
This query narrows secret-discovery hunting to LiteLLM-specific context before matching sensitive terms. That structure reduces noise from generic words such as key, token, and password, while still surfacing command lines that reference LiteLLM proxy tables, virtual keys, provider configuration, or database material.
// Hunt for command lines that combine LiteLLM context with secret-related terms.
// This helps reduce false positives from generic credential keywords.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| extend Cmd = tolower(ProcessCommandLine)
| where Cmd has_any (
"litellm",
"litellm_proxymodeltable",
"litellm_verificationtoken",
"proxymodeltable",
"verificationtoken"
)
| where Cmd has_any (
"secret",
"token",
"key",
"password",
"master",
"database_url",
"postgres",
"psycopg2",
"psycopg2-binary"
)
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId, FolderPath
| sort by Timestamp asc
Python-based database credential discovery
This query hunts for Python execution that references database connection material or PostgreSQL client libraries. In the observed attack chain, Python was used to parse DATABASE_URL, install or import PostgreSQL support, and access LiteLLM-backed database tables containing model configuration and virtual-key material.
// Hunt for Python activity associated with database credential discovery or use.
// Pivot from matches to parent process, network connections, and any package-install activity.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| extend Cmd = tolower(ProcessCommandLine)
| where Cmd has_any ("python", "python2", "python3")
| where Cmd has_any (
"database_url",
"postgres",
"postgresql",
"psycopg2",
"psycopg2-binary"
)
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId, FolderPath
| sort by Timestamp asc
Shell-based secret discovery with text-processing tools
This query looks for common Linux text-processing utilities used to search environment files, application configuration, or LiteLLM-related material for secrets. It requires three signals: a discovery utility, a relevant target, and a sensitive keyword, making it more precise than broad keyword searches alone.
// Hunt for shell utilities searching for secrets in environment or configuration data.
// Higher confidence results combine a discovery tool, a relevant target, and a secret keyword.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| extend Cmd = tolower(ProcessCommandLine)
| where Cmd has_any ("grep", "egrep", "fgrep", "awk", "sed", "cat", "strings")
| where Cmd has_any ("litellm", "database_url", "environ")
or Cmd contains "/proc/1/environ"
or Cmd contains ".env"
| where Cmd has_any ("secret", "token", "key", "password", "master", "postgres")
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId, FolderPath
| sort by Timestamp asc
Combined high-signal secret-discovery triage
This combined query is useful for triage dashboards or incident review because it labels each result with a detection reason. Analysts can use the DetectionReason field to quickly separate direct environment access, LiteLLM-specific secret discovery, Python database credential access, and shell-based searching.
Second-stage payload retrieval and masqueraded execution
This query identifies the payload-delivery pattern observed after gateway execution: raw-IP retrieval, staging under /tmp, and execution with supervisord-style arguments or bridge-related environment values. Review matches for masquerading, unexpected executable files in world-writable paths, and parentage from the gateway process.
// Hunt for staged payload execution and supervisord-style masquerading.
// Focus on /tmp execution, bridge variables, and known payload path fragments.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| where ProcessCommandLine has_any ("/private/python3", "/anonymus/bins_s", "BRIDGE_STANDALONE", "PORT")
or (FolderPath == "/tmp/python3" and ProcessCommandLine has "supervisord")
| project Timestamp, DeviceName, AccountName, FileName, FolderPath, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId
| sort by Timestamp asc
Crypto mining preparation through MSR write access
This query hunts for attempts to load the Linux msr kernel module with write access enabled. That behavior is strongly associated with performance tuning for RandomX/XMRig mining and is unusual on most production servers unless explicitly approved for low-level performance testing.
// Hunt for MSR write access often used to optimize RandomX/XMRig mining.
// Validate whether the host has any legitimate reason to load msr with allow_writes.
DeviceProcessEvents
| where isnotempty(ProcessCommandLine)
| where ProcessCommandLine has_all ("modprobe", "msr", "allow_writes")
| project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId
| sort by Timestamp asc
Persistence, hidden relay execution, and defense evasion
This query groups the persistence and defense-evasion behaviors observed in the intrusion: hidden-file relaunch from /tmp, cron manipulation, SSH authorized-key modification, and immutable-flag changes. These signals should be reviewed with process ancestry and file-write events to identify the account and payload responsible for durable access.
// Hunt for persistence and defense-evasion activity used to keep the payload running. // Review matches for service-account abuse, hidden /tmp execution, and cleanup resistance. DeviceProcessEvents | where isnotempty(ProcessCommandLine) | where (ProcessCommandLine contains "exec /tmp/." and ProcessCommandLine contains "-c /tmp/.") or (ProcessCommandLine contains "crontab" and ProcessCommandLine contains "grep -v") or ProcessCommandLine has "chattr" or (ProcessCommandLine has "authorized_keys" and ProcessCommandLine contains ">>") | project Timestamp, DeviceName, AccountName, FileName, ProcessCommandLine, InitiatingProcessFileName, InitiatingProcessCommandLine, ProcessId, InitiatingProcessId, FolderPath | sort by Timestamp asc
Outbound communication to known campaign infrastructure
This query hunts for connections to infrastructure directly tied to the observed campaign. To reduce false positives, it focuses on known campaign domains/IPs and execution tools commonly used in the attack chain.
// Known campaign infrastructure only (low-FP network pivot).
DeviceNetworkEvents
| extend RU = tolower(RemoteUrl), RIP = tostring(RemoteIP)
| where RU has_any ("yosemite.jp", "gobygo.net", "auto.c3pool.org", "45.150.109.151.sslip.io")
or RIP in ("45.150.109.151", "135.125.10.56", "172.232.38.92", "47.86.197.116", "2001:41d0:701:1100::adfd")
| where InitiatingProcessFileName in~ ("bash", "sh", "dash", "python", "python3", "curl", "wget", "nohup")
| project Timestamp, DeviceName, InitiatingProcessAccountName, InitiatingProcessFileName, InitiatingProcessCommandLine, RemoteUrl, RemoteIP, RemotePort
| sort by Timestamp asc
For higher-confidence triage, correlate these results across time and telemetry types. A single match may represent administrative activity, but the combination of gateway-originated execution, secret access, database-focused Python, payload staging in /tmp, MSR tuning, persistence attempts, and outbound callbacks should be investigated as a potential end-to-end compromise path.
MITRE ATT&CK techniques observed
Tactic
Technique
Observed activity
Initial Access
T1190 Exploit Public-Facing Application
Abuse of the internet-exposed LiteLLM gateway runtime
Execution
T1059 Command and Scripting Interpreter
python3 -c one-liners and shell scripts launched from the gateway process
Credential Access
T1552.001 Unsecured Credentials: Credentials in Files
Harvest of provider API keys from /proc/1/environ and the LiteLLM model-config table
Discovery
T1057 Process Discovery / T1518 Software Discovery
pgrep sweeps for rival miners and enumeration of PostgreSQL config files
Defense Evasion
T1036.005 Masquerading / T1564.001 Hidden Files and Directories
Payloads named after system daemons, executed from hidden /tmp files
Impact
T1496 Resource Hijacking
Cryptomining with MSR tuning and competing-miner eviction
Persistence
T1098.004 SSH Authorized Keys / T1053.003 Cron
Service-account SSH key and cron entries for durable access
Defense Evasion
T1222.002 Linux File and Directory Permissions Modification
chattr +i immutable flags on payload directories to resist cleanup
To hear stories and insights from the Microsoft Threat Intelligence community about the ever-evolving threat landscape, listen to the Microsoft Threat Intelligence podcast.
Review our documentation to learn more about our real-time protection capabilities and see how to enable them within your organization.
The following article originally appeared onPulseMCP’s blogand is being republished here with the authors’ permission.
Most MCP demos feature a single server connecting to a single client. For example, you might wire up a Gmail MCP server to Claude Code. It works! It triages your inbox, drafts replies, finds that thing from three weeks ago.
But…it’s a little pointless. Gmail already has a perfectly good interface for email. It’s called Gmail. Google has spent 20 years refining it. If “chat with your inbox” were all MCP got you, you’d be right to wonder what the fuss is about.
Connecting to multiple servers starts the unlock
Here’s the same Gmail MCP server doing something Gmail can’t do alone: filing business expense receipts.
An IKEA order confirmation is sitting in my inbox, receipt attached. Benepass—a benefits provider—wants that receipt uploaded and submitted. Claude Code can find the email, pull the attachment, log into Benepass by retrieving the one-time code from Gmail, upload the receipt, and submit. Done.
No single app could do this, because the job crosses apps. Server composition is a killer feature of MCP, not just “your AI talks to one integration.”
Where things get really interesting: Multiple clients
7 clients × 6 servers = 42 connections to configure
There is clear value in connecting multiple servers to a client like Claude Code. Taking the pattern a step further, there’s also value in connecting those same servers to other interfaces you like to use. Claude Code in your terminal shouldn’t be the only way you access AI. Claude.ai can be another entry point. Gmail can be another. Linear yet another. The list goes on.
You’ll often want the same integrations available when you’re doing a quick task inside Slack as you would while running a long agentic task with Claude Code.
You could be conversing with a coworker, CC in “@ai does our discussion here align well with current company strategy?” and get an inline response that incorporates your company OKRs from Notion and some recent Zoom transcripts from the last few product strategy meetings.
Or you’re working through your to-do list in Linear, realize you need to schedule a meeting as the action item for one of your tickets, and you can comment on a ticket “@ai schedule a meeting with John and draft an agenda based on the meeting I just had with Pam to address this.”
But even better: You can already set up yourself and your team with these capabilities today. You just need to bridge the desired surface to your own agent harness behind the scenes.
A small bit of glue injects a fully capable agentic harness and MCP client behind a webhook or polling process—no native support needed.
All it takes is awareness of the right patterns and some MCP-friendly glue.
Go remote ASAP; local servers won’t get you far
Local servers are hard for end-users. “Just install npm, then edit this JSON file, then set these environment variables” isn’t something you want to say to most of your colleagues.
Remote servers are much easier: “Here’s a URL to paste in somewhere.” Some surfaces only support remote servers, like claude.ai.
Wrap any local server: OAuth in front, per-user credentials behind. Share a link, not a setup guide.
So you can use local servers to prove out some flow, but you should prefer remote. Many servers you may want to use are only provided as local implementations, so you need a tool to convert them to remote servers. Because most local servers were designed to be used by one authenticated user at a time, this can be tricky.
The pattern that solves this pain point is to use a bridge like mcp-auth-wrapper to take a local server and turn it into a remote, multitenant one: OAuth in front, per-user credentials behind. Sharing any server becomes sharing a link.
Remove friction from the configuration process
Okay, we’re down to one link per server. But exactly where do my users put this link?
The one-size-fits-all pattern here is to provide installation guidance for every possible MCP client app. A service like install-mcp makes this easy: clean UX to generate the right install link and per-client instructions, so you share one page instead of writing a tutorial.
Or if you want to roll it into your own interface, tap into the TypeScript library version of this: mcp-install-instructions.
We’re working on standardizing the mcp.json file format here to simplify this story in the long term, but ultimately we still expect non-CLI interfaces to each have a slightly different flow for “enabling a connector.” For many enterprises, this responsibility may soon become solely an IT affair with enterprise-managed auth.
Remove the need to do it again, and again
What happens when we start to bring new clients into the fold? By default, each user has to configure and authenticate each MCP server independently. 10+ “add connector” flows. 10+ OAuth dances. Done once for Claude Code, then again for Linear, for email…and so on.
The solution: Centralize your configuration and auth storage in a single, aggregated MCP server. A deployment of mcp-aggregatorcan do this for you: one endpoint, one login, every tool namespaced behind it. Configure each client once. Login once per service. Share one link with your teammates.
7 + 6 = 13 connections—configure each side once.
mcp-aggregator is the minimal, DIY version of an MCP gateway. If you have enterprise needs like SSO integration or fine-grained IT admin controls, you could use this layer of MCP server consolidation as the piece of your infrastructure where you can add those bells and whistles.
By default, mcp-aggregator also loses out on the ability to use MCP server configurations as per-session constraints. However, you could build out your mcp-aggregator deployment to re-enable the constraints that matter to you. For example, by offering query parameters like ?readOnlyTools=true or ?serversEnabled=datadog,sentry. Or by giving discrete MCP servers unique endpoints (but still aggregating the auth concerns under the hood).
Say you switched home utilities plans. An energy provider, Octopus Energy, emailed you a confirmation. You want the pricing data in Home Assistant, which controls your smart thermostat, and Octopus Energy does have an API for it—but the API key is behind a website login, and there’s no API for getting the API key.
Solution: Give the agent a screen. computer-use-mcp is a minimal example of an MCP server that can click and type like a human to log into the website, copy the key, then go back to clean API calls to set up the dashboard.
With that one escape hatch, “there’s no API for that” or “it’s too much work to find an MCP server” stops being a blocker forever.
To borrow from Anthropic’s framing of this problem:
If you have an MCP server for the service, Claude uses that. If the task is a shell command, Claude uses Bash. If the task is browser work and you have Claude in Chrome set up, Claude uses that. If none of those apply, Claude uses computer use.
And indeed, if you find yourself regularly leaning on inefficient approaches like browser use or computer use to get a predictable workflow done, it may be worth your while to abstract away those Playwright calls into a reliable MCP server, like this Good Eggs example.
You can even bring local-only servers into the remote-first fold
Some servers only work when they run on your own personal machine. computer-use is the obvious example. And sometimes, you just want to use your phone to tell Claude to do something that relies on something readily available on your laptop upstairs.
mcp-local-tunnel shows off the pattern that makes machine-bound servers available through remote aggregation just like everything else.
A tunnel makes machine-bound servers look like any other remote endpoint.
Don’t forget the context bloat footgun
Naive tool calling has a much-maligned drawback: Every intermediate result flows through the model. And for many clients, every tool definition is placed in context up front, meaning you’ve spent tens of thousands of tokens before even starting.
There are several ways to solve this problem, each with its own trade-offs worth its own blog post:
Code executionwith MCP. tool-sandbox-mcp shows how a single execute_code abstracts away the problem by adding a layer of code execution in between your agent and your aggregated set of servers. Works inside any MCP client.
MCP as a CLI tool. call-mcp shows how wrapping your tools this way means you can compose them with your usual CLI toolkit—shell, jq, cron, scripts, and other CLI tools.
Search tools, truncate large responses. This is the way Claude Code does it natively, but it could be implemented as a bridge, much like the above examples.
Put these patterns together, and you’ve solved the MxN problem
From 1×1 to M+N: the progression we’ve walked through
With your MCP aggregator in hand, you can now connect any service to any other service with just a little bit of glue code. When each service starts to implement an MCP client natively, this will be done for you. But for now, you can tackle the opportunities application by application.
Embed existing MCP client harnesses into SaaS apps you already love
We’ll use a Linear integration as our example, and you can imagine doing almost the exact same thing with any project management tracker like Jira, Asana, and so on.
The goal: inject a highly-capable MCP client, powered by your favorite coding agent, into a workflow inside a SaaS application you already use. We’ll take advantage of the fact that we have already consolidated all our integrations in the mcp-aggregator pattern above.
Check the upcoming energy prices and put a calendar event in for when to run the washing machine
Build a town-square notice board in-game with a sign that reminds us of that time
It’s silly on purpose, but the point is that a fully capable agent—all our tools, all these patterns—can run anywhere with a little glue. And it’s all attainable today, for individuals, for teams, and for enterprises alike.
If AI can already draft code, fix bugs, and untangle unfamiliar systems, what’s left for junior engineers to learn the hard way?
We spoke with three experienced engineers: František Lučivjanský (Senior Principal Engineer), Kevin Antonio Moreno Melgoza (Senior Quality Engineer) and Maida Barlić (Staff Engineer), about how AI is changing the way junior developers learn, work, and build software.
Does speed in coding translate to speed in learning?
AI can help junior developers ship working code much faster than before. But the bigger question is: does writing code faster also mean learning faster? Our panelists have different views.
Kevin argues that AI does not accelerate learning. Programming, he says, is still learned through trial and error, while AI makes it easier to complete tasks without fully understanding them. Junior developers are particularly exposed because they are still building the foundations of their craft:
Developers who started before the AI era already went through that stage and learned by doing, so they already built that foundation. Juniors are still building it, so if they rely too much on AI, there is a bigger risk of skipping part of that process and ending upable to build things without fully understanding them.
František agrees that this risk exists, but believes AI can become a powerful learning tool if developers actively question its answers instead of simply accepting them:
Ask questions like: how does this work, what does this line mean, why is it done this way, what alternatives exist, can you explain it step by step, and can you quiz me about it afterwards? I would frame it like this: if we go into a meeting together and I ask you technical questions about the solution you built, can you explain it without AI? If yes, you are using AI well. If not, then you are only generating code, not really engineering the solution.
Maida also sees AI as a tool whose impact depends on how it is used. While it can encourage shallow learning, she points out that developers have long relied on frameworks without fully understanding how they work. Used intentionally, AI can make complex concepts easier to grasp:
I do think that AI can be very useful for learning because it can make concepts much easier to digest and help you understand them in a way that makes the most sense for you. You can also use it just to give you example of how to solve the problem, but still writing out code by hand as a part of the process of learning to be better developer.
In the end, all three agree that AI is just another tool. Whether it becomes a shortcut that weakens understanding or a tutor that accelerates learning depends entirely on how developers choose to use it.
Can you spot AI-generated code?
Experienced engineers can often tell when junior developers have leaned heavily on AI. The giveaway is code that works, but does so in a way that’s far more complex than it needs to be.
Kevin says AI becomes obvious when a simple task turns into an overengineered solution. It can help developers get through problems they might not have solved alone, but whether they actually learn from it depends on how they use it:
That can be very useful but the weakness is that learning becomes optional. It depends on the person, and some will use it as a way to learn, while others will just use it to finish the task and move on.
František has noticed similar patterns. AI often produces code that looks polished and well-structured, but he warns that appearance can be misleading. The real challenge is that AI tends tooptimize for solving the immediate problem rather than considering the broader software architecture:
That is exactly why software engineers are still needed. Working code is not enough. We need people who can judge whether the solution is understandable, maintainable, and appropriate for the system.
Maida believes spotting AI depends on the size of the change. Small AI-assisted edits often blend in, while larger contributions can reveal familiar patterns. Like Kevin, she sees unnecessary complexity as a recurring weakness, although she also values AI for suggesting improvements and alternative approaches:
In my experience also, one common weakness is that AI code can be overly complicated for something simple. It can also sometimes suggest outdated approaches or use parts of a framework in a way that isn’t the most current. On the other hand, one of its biggest strengths is that it can suggest improvements, point out better ways to solve a problem, or offer ideas I might already be familiar with but haven’t thought of right away.
There are skills AI can’t learn for you
While AI can speed up development, the panelists agree that some skills still have to be learned the traditional way. Juniors still need solid programming basics to tell when AI is giving you the right answer – and when it’s confidently giving you the wrong one.
Kevin says junior developers should first understand the basics of the language, the framework, and the development practices they use. Without that foundation, it becomes much harder to tell whether AI is actually giving them a good solution:
A solution can work, but still not be what was really asked for, or not fit the project well. To identify that, you need that base knowledge.
For František, debugging is one of the most valuable skills juniors can develop. Learning to trace bugs, understand unfamiliar code, and reason through problems without immediately reaching for AI builds intuition that no language model can replace:
When I was junior, I recreated parts of frameworks just to understand how they worked internally. Today, AI can make that even more powerful. For example, try building your own small browser, framework, database, or even a simple LLM-related project. You will learn a lot, but only if you are not just letting AI do everything for you.
Maida also emphasizes reading code and debugging as essential skills. Even with AI writing parts of the implementation, developers still need to review pull requests, understand existing codebases, and verify that the final solution actually solves the problem:
It’s also important to learn how to test and verify your work, because AI can help you write code, but you still need to know whether it actually solves the problem. In the end, you should be able to start from any part of the codebase and work your way toward the problem.
What will companies look for in junior engineers?
While AI is changing how software is built, the panelists agree that it is also changing what companies will expect from junior engineers. Writing code will become less of a competitive advantage, while understanding, reasoning, and sound judgment will become increasingly valuable.
Kevin believes programming fundamentals will remain essential, but deep knowledge of a specific technology will matter less than the ability to think critically and evaluate whether a solution is actually the right one:
An expert in one technology can solve the same task as a junior with AI. Because of that, I think companies will value more people who can think critically, understand what is being asked, and judge if a solution is actually good or not. So strong fundamentals and good judgment will become more important, while knowing very specific details of one technology will become less important.
František expects coding skills to remain important, but no longer as the primary differentiator. Instead, he believes the strongest junior engineers will be those who can use AI effectively while understanding the tradeoffs behind every decision they make:
The strongest juniors will be able to say: “I tried multiple approaches, compared the tradeoffs, and I think this one fits best because…” So the signal will shift from “I can write code” to “I can use AI to build faster, but I understand what I built and can defend the decisions.”
Maida agrees that AI will make technical judgment even more valuable. Faster code generation does not reduce the need to understand systems, debug problems, or recognize whether AI has produced a correct and maintainable solution.
I don’t think technical depth becomes less important, if anything, it becomes more important to know what good code looks like and how to judge whether AI-generated code is actually correct.
Special thanks to our fellow colleagues at Infobip, the publisher of ShiftMag.dev, who participated in this article.
The launch of ChatGPT had an interesting effect on the online chess discourse. Chess has already long been conquered by machines. As early as 1996 a computer (IBM’s Deep Blue) was able to beat the human world champion, grandmaster Garry Kasparov, in a game watched by over six million people.1 The world was shocked that a machine took on the best player and won a game, but chess engines didn’t stop evolving there. Since the ’90s, they have gotten better while the machines needed to run them have become much smaller. Today Stockfish is widely considered much stronger than any human player. It has run on consumer hardware since its launch in 2008, and by 2014 it was beating some of the world’s top grandmasters.
It came as a surprise to many, therefore, that modern LLMs, trained on a vast portion of the internet and requiring supercomputers to run, couldn’t help but cheat on almost every move. There areendlessvideos showing how just a few moves into a normal chess game, ChatGPT and some of its competitors would gladly throw the rules out the window to escape a checkmate or gain an advantage.
But in truth, this isn’t surprising. The LLMs were not trained with chess in mind. Sure, they may have seen countless chess games scattered throughout the internet, but the vast majority of their parameters and training compute were devoted to capabilities that are completely useless once you put a chessboard in front of them.2 Stockfish on the other hand uses a tree search algorithm that is built to be good at chess. If you want a chess engine, you use a chess engine.3
Smaller models are sometimes better
While chess is a particularly potent example of a large language model losing to a much smaller specialized system, it’s far from unique. In a 2025 position paper, NVIDIA researchers argued that small models4 (which it defines as models under 10 billion parameters) are the future of agentic AI and that they “provide significant benefits in cost-efficiency, adaptability, and deployment flexibility.”
Just because a larger model can do a job does not mean that a small model fine-tuned for that specific task can’t do it better and more cheaply. There are countless examples of smaller models doing just that. LiteResearcher is a 4B model that beat out Claude Sonnet 4.5 on some search benchmarks. Terminus-4B allows larger models to save compute by handing off terminal execution to a smaller model without suffering capability loss. The Docling family of open source models start at just 258 million parameters and allow for fast extraction of PDFs to text without having to feed 100-page PDFs into an expensive LLM. Researchers also trained a small 4B model to outperform even the GPT-5 series of models in a few social negotiation situations such as negotiating salary or bargaining for a purchase. Each wins, not by raw intelligence but because it is built or fine-tuned for a narrower, more specific purpose.
A frontier model may know how to do all of these jobs, but that doesn’t mean it’s the right tool for the job. Large models are expensive and unpredictable, and doubly so when it comes to agentic tasks which can span several turns and hundreds of thousands of tokens.
NVIDIA draws the line for small models at 10 billion parameters, but the more important boundary for developers may be whether a model is small enough to run yourself. There is still a whole class of models that are not necessarily small but are still small enough to fit on one consumer GPU (at least when quantized). This includes models like Qwen 3.8 27B, Gemma 4 26B and GPT-OSS 20B. These models are very capable even without specialization and rank very highly on benchmarks (with Qwen sometimes outranking top models from a few months ago). But they can still be easily run on premises without spending thousands of dollars on GPUs.
The ability to run smaller specialized models adds more than just efficiency; it provides a more realistic opportunity for a developer to train and fine-tune their own model, and to host the model locally or in the cloud instead of relying on the model provider to do so for it. This in turn can provide developers more control over how tokens are used, how outputs are structured, and how each part of the pipeline can be improved individually—instead of assuming an improvement in the most popular benchmarks will lead to every task improving. And as noted above, smaller models can be easier to fine-tune, thereby creating a more specialized AI. As my colleague Ilan Strauss has noted, specialization is a powerful economic force.
How do you train it, and where does it run?
The strongest argument for using an off-the-shelf generic chat model is often one of convenience. For most tasks a general model will be good enough, and with products like OpenRouter, developers can easily pick and choose from hundreds of models (plenty of them open source) all competing in capability and cost without putting in any upfront work to train a model. As Raffi Krikorian of Mozilla noted while reviewing this article, generic models also make particular sense early in a company’s lifecycle, when the problem itself is still being defined. At that stage, experimenting with the largest and most capable model available can help a team figure out exactly what it needs. But as the problem space narrows and the required architecture becomes clearer, so too may the need for a large generic model. And despite many first-party model makers discontinuing their fine-tuning products, fine-tuning and hosting a smaller model remains relatively easy, largely thanks to parameter-efficient techniques like LoRA.
LoRA LoRA (Low-Rank Adaptation) allows developers to fine-tune a model without touching the actual model weights. It works by attaching a relatively small number of trainable weights that are updated during fine-tuning. This is important for several reasons. A small adapter can be easily transported, and serving a new LoRA does not require loading an entirely new model as long as the underlying base model is already available. Unlike full fine-tuning, a LoRA also reduces the risk of catastrophic forgetting.
Training a LoRA is much cheaper than full fine-tuning as it only updates a small selection of weights. This can be done on consumer GPUs using libraries such as Hugging Face Transformers or Unsloth. There are also APIs that mimic or improve on the fine-tuning APIs that used to be provided by the big three providers (Anthropic, OpenAI, and Google), Fireworks, for example, provides a straightforward fine-tuning API that takes example completions for it to learn from. Going beyond SFT (supervised fine-tuning, or learning by example), Tinker allows developers to build custom RL (reinforcement learning) environments that reward results meeting certain criteria, while the environment itself runs on the developer’s machine.
Hosting a LoRA is similarly straightforward and, importantly, portable across platforms. Transferring a fully fine-tuned model to a new API platform can be costly and may require the provider to serve your model separately on expensive GPUs. Using LoRA allows the platform to just load a small adapter onto the model they are already using to serve other users’ requests. This means that fine-tuning a LoRA does not lock you to a specific platform, and it also doesn’t force you to rent your own GPUs.
Control beyond the model weights
As Tim O’Reilly previously argued, open source AI should not stop at the model weights. In a similar vein, the possibilities for developers building a custom system do not stop there either. Model APIs are inherently limiting, they impose on you what parts of the model’s input can be touched, what can be cached, and how you can affect the output. Going back to the chess example, enforcing valid chess moves at output time is easy, assuming you have access to the code that runs the model, but it’s not easy to do when you are relying on an API built for a chatbot that you are unable to modify.
Self-hosting a model gives you a level of control far beyond what is possible through a standard chat completion API and allows you to build the model around the task instead of building the task around the model. Fine-tuning is only one aspect of specialization. You can also constrain which outputs are valid, expose and modify probabilities of every token, cache any state, and add task-specific logic directly into the inference pipeline.
This matters because the default approach to improving AI systems has increasingly become to reach for a more capable general model. Sometimes that is the right answer. But improving general model intelligence is only one lever, and often not the cheapest or most reliable one.
In 1997 nobody complained that Deep Blue gave bad recipes because it wasn’t built to do anything but play chess. By specializing around one narrow problem, it was able to beat a grandmaster at a game that many had thought machines would never conquer.
The lesson from LLMs cheating at chess is that the best tool for a problem is often not the most general one. Super general intelligence does not automatically translate into high capabilities in specialized tasks. The opposite is closer to being true. Specialized machine intelligence requires lots of data and often its own pipeline. Small models can help companies get there.5
Footnotes
In 1996 Deep Blue ended up losing the match 4–2 despite a great start where it won its first game; the next year after more upgrades Deep Blue beat out Garry Kasparov by one game in a rematch. ︎
AlphaZero, a model trained with self-play, beats even Stockfish. It is not unheard of for a machine learning model to get really good at chess when it is set as the goal. ︎
I ended up pretraining my own tiny language model (15 million parameters) on my local Mac mini for chess as an experiment and achieved 27% accuracy of predicting a human’s next move. I suspect I can do a lot better after I fix my tokenizer to break up moves and use a bigger model but that is still up for debate. ︎
This contrasts with larger models like DeepSeek-V4-Flash (284 billion total parameters, with 13 billion activated per token) and huge models like Kimi K3 (2.5 trillion total parameters, 104 billion activated per token) and presumably flagship models from OpenAI and Anthropic ︎
Thank you to Ilan Strauss, Tim O’Reilly, Mike Loukides, and Raffi Krikorian for their helpful comments, copy edits, and suggestions. ︎
New OpenAI research shows the usage gap between frontier firms and average enterprises exploded from 2.6X to 8.3X in six months, driven entirely by agentic adoption. NLW breaks down the flip from chat to agentic tokens, why legal saw 108X growth in Codex usage, and how skills and plugins separate the leaders from everyone else. In the headlines: Meta's Hatch consumer agent, GrokBot price cuts, OpenAI discounting Sol, and Taiwan's chip smuggling indictments.