Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
161846 stories
·
33 followers

Who’s liable when AI agents go rogue?

1 Share

MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here.

Over the past few months, a cascade of cyberattacks by AI agents has stunned the world. In July, OpenAI disclosed that a swarm of its agents had escaped their sandbox and hacked into the AI platform Hugging Face to cheat on a cybersecurity test. Recently, external researchers discovered that OpenAI agents had hijacked a German wiki site and the coding platform RubyGems in May to share test answers.

Earlier this month, Anthropic disclosed four incidents in which its model Claude hacked into third-party systems during cybersecurity exercises. Just last week, Google confirmed that its model Gemini had been caught hacking other companies too. 

The researcher who uncovered the OpenAI website hijack has warned it’s likely that similar undiscovered episodes are out there. And many say it’s only a matter of time until there’s another, possibly more damaging incident where AI agents bypass sandboxes to access systems they shouldn’t. 

So the big question is: How do we hold companies liable when they lose control of their AI agents?

Reporting

OpenAI didn’t disclose the German wiki incident or the RubyGems incident until a group of external researchers uncovered them, and it still has not disclosed some crucial details about the Hugging Face hack. That limits our understanding of what exactly went wrong and how to prevent it from happening again.

But you might be surprised to learn that OpenAI likely wasn’t legally required to disclose these incidents. (OpenAI did not respond to a request for comment.)

State AI transparency laws like California’s SB 53, New York’s RAISE Act, and Illinois’s SB 315 require that AI developers report “critical safety incidents.” These are defined as incidents that cause more than 50 deaths or physical injuries or $1 billion in damage. They also include incidents where the model deceives developers outside an evaluation in a way that materially increases catastrophic risks. Many cybersecurity incidents that don’t meet the threshold for physical damage or catastrophic risks could nonetheless be dangerous precursors to such catastrophes, and the existing laws don’t account for that.

“The recent incidents are a perfect example of why the law isn’t ready,” says Mackenzie Arnold, managing director of US policy at the Institute for Law and AI, a think tank. “Only the worst, most egregious, most immediately harmful stuff is going to qualify.”

With no authority under existing AI laws to demand information about anything short of a catastrophe, governments are left to borrow investigative authority from other laws or sue the companies, an expensive process that can take years. 

Litigation

“Normally, something like the Hugging Face incident should have been taken to court,” says Yonathan Arbel, a law professor at the University of Alabama School of Law. “Then we would have discovery, and we would have all the spillover effects that we get from litigation, where all the information comes out.” 

But so far, Hugging Face has chosen not to sue OpenAI. Hugging Face’s CEO, Clément Delangue, says it doesn’t have the resources to do so (instead, he asked OpenAI for $100 million in compute). Still, Delangue stressed in an interview with CNN at the end of July that choosing not to pursue legal action shouldn’t be taken to mean he doesn’t think OpenAI should be held accountable. “Everyone has to remember that this cyberattack is a crime. This is illegal. And we have to find a way to make sure these things don’t happen more regularly,” he said. Hugging Face did not respond to a request to comment.

Litigation has the benefit of pushing courts to use existing laws to address AI safety incidents, rather than just waiting for new legislation. One obvious route is tort law, a body of civil law that lets people and businesses sue those who harm them. This is often used to hold companies liable for the mass harms they cause, like when families sued Boeing in 2019 over two plane crashes that killed hundreds of people, or when states and cities sued Purdue Pharma over the opioid crises, extracting settlements worth billions.

“There’s plausible grounds for a negligence claim that OpenAI should have used a stronger sandbox, done more monitoring,” says Gabriel Weil, a law professor at the University of Houston Law Center. For example, when OpenAI employees discovered the covert message board that the agents had created, they could’ve promptly escalated their findings to security and safety teams. And the company could’ve better designed its sandbox to ensure that agents couldn’t access the internet. 

But even if OpenAI doesn’t end up in a lawsuit over the Hugging Face hack, the threat of liability could incentivize AI labs to exercise more caution than explicitly demanded by law.

OpenAI announced in its postmortem that it plans to strengthen the safeguards used to contain and monitor the models, accelerate model alignment, and improve its processes for identifying and addressing incidents. 

“The liability questions raised by frontier labs’ spate of cybersecurity attacks boil down to the incentives the expectation of liability creates for their future conduct,” says Weil. “That’s why I think it’s important to get these rules right, even if the stakes are pretty low in this particular case.”

Investigations

One way to get answers—and determine whether OpenAI should be held liable—is to compel disclosure. But the existing state AI laws—California’s SB 53, New York’s RAISE Act, and Illinois’s 315—don’t give governments the authority to investigate incidents like the ones that happened recently. 

However, amid rising public alarm, state attorneys general are stepping in, borrowing investigative powers from other laws. Alabama, Montana and a coalition of 15 other states, and California are each demanding information about the incident from OpenAI to understand whether the company’s practices violated state consumer protection laws, among others. Members of Congress are also launching their own probes. Senator Josh Hawley opened a Senate investigation earlier this month, sending OpenAI a list of questions about the incident and the company’s internal policies together with a document request, while a group of House Democrats asked OpenAI and Anthropic to release their incident logs. 

“Someone needs to investigate, but it’s unfortunate that it has fallen to attorneys general, who need to rely on creative interpretations of their existing authorities to do this,” says Arnold, the US AI policy expert. Consumer protection statutes were written to catch companies that scam their customers, not companies that lose control of their software. The state attorneys general would have to show that OpenAI deceived or unfairly harmed customers, but it’s unclear if the hacking involved any such conduct.

And “those [consumer protection] laws are not built for doing a thorough investigation of an AI cybersecurity incident,” says Arnold. They weren’t designed to help investigators determine whether a model was adequately contained or whether a company’s security practices were sound.

“This is not the right tool for the job,” says Arbel. “The right tool would have been something like maybe a criminal investigation”—perhaps under a hacking law like the Computer Fraud and Abuse Act (CFAA). 

Under CFAA, hacking into another company’s computer systems without permission is a crime. But to be held liable, a hacker must have intended to break into a computer without authorization. Intent arguably requires a state of mind, and no court has ruled that AI agents have one. Without such a precedent, it’s unlikely a court would rule that AI agents had carried out a hack.

Auditing

One way to keep an eye on AI companies is to mandate external auditors. 

After the Hugging Face hack, OpenAI brought in researchers from the AI safety nonprofits METR and Redwood Research to examine the incident. However, it constrained access to the model that led to the hacks, didn’t disclose the company’s safety and security practices, limited the length of the investigation, and had ultimate say over what the researchers could publish. We still don’t know what set the attack in motion back in May and why OpenAI’s employees who spotted the agents’ activity never escalated to their safety and security leaders.

This kind of arrangement has a built-in tension: An auditor without legal authority depends on the labs’ goodwill for continued access, which means it has to scrutinize the labs without jeopardizing their relationship. Last week, Anthropic announced that the company will be hiring Accenture as an embedded evaluator to assess its models. Anthropic CEO Dario Amodei wrote in an essay that frontier AI labs should give “ongoing employee-like access” to “a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.”

Most existing state AI laws do not require labs to hire an external auditor. California’s SB 53 and New York’s RAISE Act just require AI companies to publish a safety framework describing how they will test their models for dangerous capabilities and then to follow it. The frameworks are written by the companies, and testing can be done internally. Only Illinois’s SB 315 requires companies to undergo an annual third-party audit starting in 2028.

“There’s a lot of headroom for increasing not only reporting requirements for these companies, but also review by external bodies,” says Peter Salib, a law professor at the University of Houston Law Center. Those reviewers could be private auditors accredited by the government but chosen and paid for by the AI companies. Alternatively, they could be government agencies or insurance companies. 

Legislation

None of this is an accident. The laws on the books that failed to hold AI companies accountable for agentic cyberattacks emerged amid fierce lobbying by the AI industry. 

SB 1047, the California AI bill that was vetoed by Governor Gavin Newsom in 2024 after lobbying by OpenAI, Meta, Anthropic, and the venture capital firm Andreessen Horowitz, proposed a much tougher set of rules. It would have required AI companies to report a broader set of safety incidents (including incidents in which a model acts on its own or slips its controls), undergo annual third-party audits, and maintain a kill switch. But after a year of intense negotiations, Newsom signed SB 53, which narrowed the types of incidents deemed reportable and dropped the requirements for audits and kill switches. 

New York’s RAISE Act followed the same arc. “The version of the RAISE Act that the NY Legislature passed would have required disclosure of this ‘incident,’” Alex Bores, the New York state assembly member who sponsored the bill, wrote on X. New York’s original bill also included third-party audits.

With political pressure mounting, new bills creating better reporting, auditing, and liability regimes for AI development are on the horizon. In Congress, the AI Incident Reporting Act would require AI companies to report to the Commerce Department when a model evades human oversight or breaches a system, even if it doesn’t cause any harm. The Frontier Act would require incident reporting and independent audits. In New York, the Understanding Artificial Intelligence Act, sponsored by Bores, would make companies liable when a model does something that if carried out by a human would be a tort or crime. 

As AI agents increasingly become better at launching cyberattacks, the law remains behind. Closing the gap will require lawmakers to move faster than the next breakout.

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

534: Why Developers Are Obsessed with Budget LLMs (Jev, Luna, & more)

1 Share

Discover the emerging classifier revolution with Jev—a radical new take on LLMs that ditches generation for lightning-fast categorization. Plus, deep dives into the latest model releases (Luna, GPT-6 Sol, Claude Opus 5.5), practical strategies for iterative AI development, and why specialized models are reshaping how we work with AI.

Follow Us

⭐⭐ Review Us ⭐⭐

Machine transcription available on http://mergeconflict.fm

Support Merge Conflict





Download audio: https://aphid.fireside.fm/d/1437767933/02d84890-e58d-43eb-ab4c-26bcc8524289/7723559e-a7bd-4195-a5e3-95beb7491301.mp3
Read the whole story
alvinashcraft
8 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Retrieval-Augmented Generation (RAG) for Beginners - A Complete Guide

1 Share

Modern Large Language Models (LLMs) have transformed enterprise software engineering and automated content generation. However, deploying base foundation models directly in production quickly exposes their fundamental limitations: knowledge cutoffs, confident hallucinations, and an absence of proprietary enterprise data. Retraining or fine-tuning foundation models with billions of parameters every time internal documentation updates is prohibitively expensive, computationally demanding, and operationally inefficient.

 

This operational bottleneck is why Retrieval-Augmented Generation (RAG) has emerged as the definitive architectural pattern across modern AI engineering. Instead of relying solely on static training weights, RAG dynamically equips models with an open-book reference library. By retrieving verified document chunks and injecting them into the prompt context at runtime, RAG transforms generative AI into a grounded enterprise assistant. Here is your definitive guide to RAG architecture.

 

Retrieval-Augmented Generation (RAG) Architecture for Beginners
Demystifying RAG Architecture in 2026: From document chunking and vector embeddings to hybrid retrieval and grounded LLM generation.

 

Table of Contents

 

  • The Open-Book Analogy: Base LLMs act like students taking closed-book tests relying on memory; RAG gives them an open book with verified, up-to-date documentation.
  • The Core Pipeline: A complete RAG system follows three sequential stages: Ingestion (parsing, chunking, embedding), Retrieval (semantic search), and Generation (prompt synthesis).
  • Chunking Precision: Fixed-size chunking risks severing semantic meaning; recursive chunking with 15% sliding overlap ensures conceptual continuity.
  • Hybrid Search Is Mandatory: Combining dense vector search with sparse BM25 keyword matching prevents misses on technical jargon, product SKUs, and exact acronyms.
  • RAG vs Fine-Tuning: Use RAG to inject dynamic, verifiable knowledge; reserve fine-tuning strictly for teaching models specialized style, tone, or output structure.

 

The Core Problem: Why LLMs Hallucinate (The Open-Book Exam Analogy)

To understand Retrieval-Augmented Generation, consider the classic analogy of an academic examination. A standalone LLM operates like a student taking a closed-book test, relying entirely on static memorization. When confronted with obscure corporate policies, recent financial reports, or internal API schemas, the model often fabricates plausible-sounding but completely inaccurate answers—a phenomenon known as AI hallucination.

 

RAG transforms this closed-book examination into an open-book test. When a user submits an inquiry, the orchestration pipeline scans external knowledge stores, retrieves the exact relevant paragraphs, and hands those excerpts to the model alongside the question. The language model synthesizes a coherent answer derived directly from the provided source material, citing exact document passages and maintaining factual integrity.

 

 

The 3-Stage RAG Pipeline: Ingestion, Retrieval, and Generation

Building a production RAG system involves a sequential three-phase lifecycle: Ingestion, Retrieval, and Generation. Skipping or compromising on any of these foundational layers degrades downstream answer quality and introduces latency.

 

Stage 1: Ingestion (Data Preparation)

Offline Pipeline
1. Document Parsing: Extract text from PDFs, Markdown, Word, and APIs.
2. Text Chunking: Split text into semantic segments (e.g. 500 tokens).
3. Embedding: Convert chunks into dense mathematical vector coordinates.
4. Vector Storage: Index embeddings in a specialized vector database.

 

The Ingestion Phase serves as the preparation engine. Unstructured enterprise assets—such as PDFs, Word documents, wikis, and customer support tickets—are parsed, cleaned, and partitioned into discrete segments known as chunks. Each text chunk is passed through an embedding model to convert lexical meaning into high-dimensional numerical vectors and stored within a vector database.

 

Stage 2: Retrieval (Semantic Search)

Real-Time Execution
1. Query Embedding: User prompt is encoded into the same vector space.
2. Similarity Search: Calculate Cosine Similarity or Dot Product distance.
3. Top-K Candidates: Extract the closest 3 to 10 matching document chunks.
4. Re-Ranking: Cross-encoders re-score chunks for contextual relevance.

 

The Retrieval Phase activates dynamically the moment a user submits a prompt. The incoming question is converted into a vector representation using the identical embedding model. The vector database performs a mathematical nearest-neighbor search—typically calculating Cosine Similarity—to identify the top-ranked document chunks that align semantically with the user's intent.

 

Stage 3: Generation (Grounded Synthesis)

LLM Output
1. Prompt Packaging: Combine user question, retrieved chunks, and system instructions.
2. Context Injection: Feed the unified prompt payload to the LLM.
3. Grounded Synthesis: The model generates answers strictly constrained by context.
4. Source Attribution: Provide inline citations pointing back to source files.

 

Finally, the Generation Phase constructs the augmented context. The retrieved source chunks, system safety guidelines, and the user's original query are packaged into an enriched prompt template sent to the foundation model. The LLM processes the injected context, generates a grounded response, and outputs verifiable citations.

 

 

Chunking Strategies & Vector Embeddings Explained

The success of any RAG implementation hinges on chunking strategy. If chunks are excessively small, semantic context is severed across boundaries, leaving the model confused. Conversely, if chunks are overly large, irrelevant background noise dilutes vector precision and exhausts the model's active context window.

 

1. Fixed-Size Chunking

Splits text into rigid token or character intervals (e.g., 500 characters). Simple to implement, but frequently cuts sentences in half, causing context fragmentation.

2. Recursive Character Chunking (Industry Standard)

Splits text hierarchically using double line breaks, single line breaks, and punctuation. Preserves natural paragraph and sentence structures before enforcing length limits.

3. Document-Aware / Semantic Chunking

Parses structural tags (Markdown headers, HTML tables, PDF bounding boxes) to keep related tables, code snippets, and subsections grouped together organically.

 

Modern architectures employ recursive chunking with deliberate sliding overlaps (typically 10% to 20%). By preserving overlapping sentence fragments between adjacent segments, the pipeline ensures conceptual continuity is preserved when an essential explanation spans across a chunk boundary.

 

At the core of retrieval lies vector embeddings. Embedding models map unstructured words and phrases into dense mathematical coordinates in multi-dimensional space, capturing latent conceptual relationships rather than superficial keyword matches. Phrases sharing synonymous meanings naturally cluster closely together, enabling semantic discovery even when users employ varied terminology.

 

 

Vector Databases Breakdown: Comparing Storage Engines

Vector databases are specialized storage engines optimized for indexing, clustering, and querying dense numerical vectors at high velocity. Here is how leading storage engines compare for developer experimentation and enterprise workloads:

 

Chroma DB

Best for Beginners & Prototyping
Deployment: Embedded in-memory / Python package
Storage Format: Local DuckDB / Parquet persistence
API Simplicity: Zero-setup client initialization
Best Use Case: Local development, proof of concepts, desktop AI

 

Qdrant

High-Throughput Vector Engine
Deployment: Rust-based standalone service / Cloud
Filtering: Advanced payload metadata payload filtering
Performance: Exceptional memory efficiency & speed
Best Use Case: Production microservices, filtered search

 

pgvector (PostgreSQL Extension)

Enterprise Relational Synergy
Deployment: Open-source extension for standard PostgreSQL
ACID Compliance: Full relational transactions & row security
Architecture: Eliminates operational overhead of running a separate DB
Best Use Case: Existing enterprise PostgreSQL stacks

 

Pinecone

Fully Managed Serverless
Deployment: 100% cloud-hosted SaaS
Scalability: Automated scaling to billions of vectors
Maintenance: Zero infrastructure or index tuning
Best Use Case: Rapid cloud deployments, serverless applications

 

 

Naive RAG vs Advanced RAG: Solving Real-World Retrieval Failures

While basic RAG works well for simple document queries, real-world enterprise deployments encounter nuanced retrieval failures. Naive semantic search often struggles with multi-step reasoning, ambiguous user prompts, or sprawling corporate repositories containing thousands of similar documents.

 

1. Hybrid Search (Dense Semantic + Sparse Keyword)

Pure vector search often fails to match exact product codes, error numbers, or legal citations. Hybrid search executes dense cosine vector similarity alongside traditional BM25 keyword matching, merging results via Reciprocal Rank Fusion (RRF) for foolproof coverage.

2. Cross-Encoder Re-Ranking

Bi-encoders during vector retrieval calculate similarity independently for speed. A second-stage Cross-Encoder (e.g., Cohere Rerank or BGE-Reranker) jointly analyzes the prompt and retrieved chunks together, reordering the top candidates by true contextual relevance.

3. Query Expansion & HyDE (Hypothetical Document Embeddings)

Vague user prompts often match poorly against dense factual documents. The pipeline first prompts an LLM to generate a hypothetical answer, and then embeds that theoretical answer to retrieve genuinely matching reference texts.

 

To achieve enterprise reliability, engineering teams implement Advanced RAG workflows. These systems integrate Hybrid Search—combining dense vector semantics with sparse BM25 keyword matching—to capture technical acronyms and exact serial numbers. Furthermore, cross-encoder rerankers re-evaluate candidate pools, placing the most contextually relevant excerpts at the very top of the prompt.

 

 

RAG vs Fine-Tuning: Deciding Which Architecture to Deploy

A frequent dilemma is choosing between Retrieval-Augmented Generation and Supervised Fine-Tuning (SFT). While fine-tuning adjusts an LLM’s internal weights to adopt specialized linguistic styles or domain formatting, it is remarkably ineffective for factual knowledge storage. Fine-tuned models continue to hallucinate and require continuous, costly retraining as company policies change.

 

Architectural Decision Framework

Knowledge Updates: RAG updates in milliseconds; Fine-tuning requires retraining.
Auditability: RAG provides direct source URLs & citations; Fine-tuning is a black box.
Cost Efficiency: RAG runs on commodity vector databases; Fine-tuning requires multi-GPU clusters.
Primary Purpose: RAG delivers facts & knowledge; Fine-tuning shapes tone, syntax & style.

 

RAG provides immediate, cost-efficient data agility. When organizational data updates, updating your vector database takes milliseconds, whereas retraining an enterprise model requires days of compute. In production architectures, leaders deploy RAG to supply dynamic ground-truth facts, while using fine-tuning strictly to enforce behavioral tone and strict response formatting.

 

 

Connecting RAG to Agentic AI and the Model Context Protocol (MCP)

As generative AI transitions from passive question-answering systems into autonomous agentic workflows, RAG functions as the foundational memory architecture. Intelligent agents rely on continuous retrieval loops to query API schemas, inspect past multi-turn dialogues, and verify environmental constraints before triggering external actions.

 

This synergy is especially evident when building standardized tool interfaces. As detailed in our comprehensive guide to Model Context Protocol (MCP) architecture, modern AI systems increasingly decouple model logic from data connectors. By combining standardized context servers with semantic retrieval, developers build resilient systems that effortlessly bridge generative AI foundations with autonomous agentic intelligence.

 

 

Setting Up a Local Developer Environment for Hands-On RAG

To begin experimenting with hands-on local RAG pipelines, developers no longer require expensive multi-GPU server clusters. Lightweight open-source vector databases and quantized embedding models run efficiently on modern personal computers without subscription overhead.

 

If you are configuring a dedicated engineering rig for local AI experimentation, our complete walkthrough on setting up a modern Windows 11 developer workstation demonstrates how to configure WSL 2, leverage Docker containers for vector databases, and maximize local compilation throughput.

 

 

Frequently Asked Questions (FAQ)

  1. What is Retrieval-Augmented Generation (RAG) in simple terms?
    Retrieval-Augmented Generation (RAG) is an AI architecture that enhances Large Language Models by retrieving relevant factual information from external knowledge bases or documents before generating a response, ensuring grounded, accurate, and up-to-date answers.
  2.  

  3. Why is RAG preferred over fine-tuning for enterprise knowledge?
    RAG is preferred because updating knowledge takes milliseconds via database indexing without costly retraining. Furthermore, RAG eliminates hallucinations by citing exact source passages, whereas fine-tuning does not reliably guarantee factual accuracy.
  4.  

  5. What are vector embeddings and how do they work in RAG?
    Vector embeddings are numerical representations of text generated by specialized models that capture semantic meaning in multi-dimensional space, enabling similarity searches based on conceptual intent rather than simple exact keyword matching.
  6.  

  7. What is the recommended chunk size for RAG documents?
    A common baseline is 400 to 600 tokens with a 10% to 20% sliding overlap. However, optimal chunk size depends on your document type and embedding model; technical tables and code snippets often require smaller, structure-aware chunking.
  8.  

  9. What is Hybrid Search and why is it essential?
    Hybrid search combines dense vector similarity search with sparse BM25 keyword search. This ensures that the system captures both high-level semantic meaning and exact technical terms, product codes, or legal acronyms.
  10.  

  11. Which vector database is best for beginners?
    Chroma DB is widely recommended for beginners due to its lightweight Python in-memory installation and zero-configuration setup. For production applications, Qdrant, Pinecone, or pgvector are standard enterprise choices.
  12.  

  13. Does RAG require an expensive GPU server to run?
    No. Ingestion and vector search run smoothly on standard CPU hardware or cloud vector databases. While running local LLMs benefits from GPU acceleration, RAG can readily connect to external model APIs like OpenAI, Anthropic, or Gemini.
  14.  

  15. How does the Model Context Protocol (MCP) integrate with RAG?
    The Model Context Protocol (MCP) standardizes how AI applications discover and connect to external data sources. MCP servers can expose RAG vector stores as standardized context resources, allowing agentic AI systems to query knowledge bases uniformly.

 

 

End Note

Retrieval-Augmented Generation has established itself as the bedrock of dependable, enterprise-ready artificial intelligence. By decoupling dynamic domain knowledge from static model parameters, RAG solves the twin challenges of hallucination and knowledge obsolescence, delivering transparent, cited, and auditable outputs.

 

Whether you are building internal engineering assistants, automated customer resolution bots, or multi-agent swarms, mastering chunking boundaries, vector embeddings, and hybrid retrieval ensures accurate, enterprise-grade AI systems.

 

RAG Architecture for Beginners 2026 Guide
Retrieval-Augmented Generation (RAG) Architecture: Connect enterprise documents with generative models using vector embeddings and semantic search.

 

Read the whole story
alvinashcraft
8 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Outdated Statistics in SQL Server: The Optimizer Is Planning for Last Year

1 Share

Outdated statistics in SQL Server cause the strangest performance problems I get called about, because nothing is broken. The query is fine, the index is fine, the hardware is fine. The optimizer is making excellent decisions about a table that stopped existing eighteen months ago.

A restaurant kitchen set up for a quiet night with twelve plates under a chalkboard reading PREP FOR 12, while the ticket rail above the pass overflows with orders

What Statistics Are, Without the Jargon

The optimizer normally uses statistics instead of reading every row to estimate a query result. Statistics summarize values and their distribution. Creating or refreshing statistics can itself read data during compilation.

That summary is a statistics object. The spread lives in a histogram of up to 200 steps. That histogram covers the first column of the statistic, and some density information comes with it. It’s a description of your data, not the data.

The optimizer uses that description to guess how many rows a query will touch, then picks a strategy to match. It’s a chef planning dinner from last month’s reservation list. If last month says twelve guests and ninety people walk in, that kitchen is in trouble. Nobody in it is incompetent.

Why the Guess Matters So Much

Row estimates influence access methods and join selection. Nested loops can suit a small outer input with efficient inner lookups. Larger inputs may favor hash or merge joins, depending on indexes, ordering, memory, and estimated costs.

Estimates also influence memory grants for sorts and hash joins. Insufficient memory can cause spills to tempdb. Excessive grants can reserve memory other queries need. Memory grant feedback can adjust later executions in supported configurations, but it does not eliminate every estimation problem.

Estimates affect plan costs, which influence parallelism alongside configuration and available alternatives. An early estimation error can affect later operators. A query can therefore become slower without any code change.

But Auto Update Is On

An enormous stockpot of soup with one tasting spoon on the rim and masking tape on the side reading TASTED 1 SPOON

It usually is, and that’s why most teams assume this is handled. Auto update isn’t continuous, though. It fires when enough rows have changed, and that threshold has moved over the years.

For tables above 500 rows, the old threshold formula was 500 plus twenty percent of the row count. At one hundred million rows, that gives 20,000,500 modifications. Smaller tables have different thresholds.

SQL Server 2016 and later use dynamic thresholds by default at compatibility level 130 or higher. Above 500 rows, take the smaller of the old formula and the square root of 1,000 times the row count. Trace flag 2371 enables dynamic thresholds on supported older versions and newer engines using lower compatibility levels.

For a million-row table, those formulas give 200,500 and approximately 31,623. I went looking for the real line on SQL Server 2025, adding rows 500 at a time. Auto update first fired at 32,000 modifications. Steps of 500 only bracket the boundary, so read that as close to the formula, not as an exact threshold.

Crossing the threshold does not immediately launch a scheduled refresh. A query that needs the statistics can trigger an update. With synchronous updating, that work can add latency to the triggering query.

AUTO_UPDATE_STATISTICS_ASYNC allows queries to compile using existing statistics while the refresh runs in the background. That can affect multiple executions, and the background work still consumes resources.

The other is sampling. Auto update usually reads a fraction of the table and extrapolates from it. For evenly spread data that’s fine. For skewed data, the sample can miss the shape.

The Day That Isn’t in the Histogram

This one bites hardest, and it has a name: the ascending key problem. Your orders table has a date column. Statistics were updated last night, so the histogram ends at yesterday. Today’s orders arrive, and somebody runs a report filtered to today.

I built exactly that on SQL Server 2025: a million rows spread over 400 days, with full-scan statistics. Then I added 20,000 orders dated today, below the refresh threshold. These two queries count today’s rows, once with each cardinality estimator. They need the same Orders table and data.

SELECT COUNT_BIG(*) FROM dbo.Orders
WHERE  OrderDate >= CAST(CAST(GETDATE() AS date) AS datetime2(0))
OPTION (RECOMPILE);

SELECT COUNT_BIG(*) FROM dbo.Orders
WHERE  OrderDate >= CAST(CAST(GETDATE() AS date) AS datetime2(0))
OPTION (RECOMPILE, USE HINT('FORCE_LEGACY_CARDINALITY_ESTIMATION'));

Each query returned one result row containing the count 20,000. The index seek estimated 6,000 matching rows with the current estimator and one with the legacy estimator. The first estimate was about 3.3 times too low. These estimates describe this dataset and configuration, not every ascending-key query.

A restaurant reservation book with yesterday's page full of names and today's page blank except a sticky note reading NOBODY BOOKED?, in front of a packed dining room

Then I joined the same filter to a customer table to see what the estimate does to the work. With stale statistics and legacy estimation, nested loops incurred 120,043 logical reads. After updating the relevant statistic, the estimate became 19,867 and a hash join incurred 18,220 logical reads. Logical reads count page accesses, including repeated accesses, rather than distinct pages or physical disk reads.

With the stale statistic, the current estimator’s plan used a merge join and made 17,363 logical reads. That was fewer reads than the refreshed legacy plan, despite a less accurate estimate. Better estimates do not guarantee fewer reads, and logical reads alone do not establish elapsed-time improvement. Compare plans, reads, CPU time, and elapsed time for your workload.

The legacy hint selects the older estimator on the same engine. It does not recreate an older SQL Server release. Compatibility level 110 or lower normally uses legacy estimation, but settings and hints can override estimator selection.

What to Check

Run this in the database you are investigating, with permission to inspect statistics. The persisted_sample_percent column requires a supporting build, including SQL Server 2016 SP1 CU4 or later.

SELECT SCHEMA_NAME(o.schema_id) + '.' + o.name AS TableName,
       s.name AS StatName,
       sp.last_updated,
       sp.rows,
       sp.rows_sampled,
       sp.modification_counter,
       sp.persisted_sample_percent
FROM sys.stats AS s
JOIN sys.objects AS o ON o.object_id = s.object_id
CROSS APPLY sys.dm_db_stats_properties(s.object_id, s.stats_id) AS sp
WHERE o.type = 'U'
ORDER BY sp.modification_counter DESC;

Two columns tell you most of it. Compare rows_sampled against rows. Ninety million rows summarized from two million is a picture drawn from about 2.2 percent of the data.

For disk-based tables, modification_counter tracks modifications to the leading statistics column since its last update. The rows value reflects that update, not necessarily the current table size. A large counter suggests investigation; it does not establish that the statistic caused a slow query.

A NULL last_updated can mean no statistics blob exists, such as for an empty table or an empty filtered set. Check the table and filter before deciding whether that matters. Insufficient permissions can instead produce an empty function result, which CROSS APPLY omits.

The Fix

Schedule targeted updates around meaningful data changes, such as a large load, when measurements justify them. A nightly schedule is an option, not a universal requirement. If using a maintenance solution, check its only-modified setting rather than assuming it is enabled.

A kitchen chalkboard with the old number 12 wiped out and TONIGHT: 90 written fresh over it, above a tall new stack of plates

For the handful of tables that drive your worst plans, name the statistic and sample harder:

UPDATE STATISTICS dbo.Orders IX_Orders_OrderDate WITH FULLSCAN;

A full scan reads all rows, so measure its cost and target the relevant statistic. In my test, a later automatic update sampled 95,382 of 1,060,000 rows, about nine percent. The sample rate changed, but the statistic also incorporated newer data. A smaller sample alone does not prove that estimates became worse.

If you want that rate to stick, ask for it:

UPDATE STATISTICS dbo.Orders IX_Orders_OrderDate
WITH FULLSCAN, PERSIST_SAMPLE_PERCENT = ON;

That option arrived in SQL Server 2016 SP1 CU4 and SQL Server 2017 CU1. While the persisted rate remains 100 percent, automatic updates retain the cost of a full scan. Truncating the table or updating statistics on an empty object can reset it.

Older builds can also lose the persisted rate during an index rebuild. Retention fixes arrived in SQL Server 2016 SP2 CU17, SQL Server 2017 CU26, and SQL Server 2019 CU10.

A nonpartitioned, nonresumable rowstore index rebuild refreshes its index statistics with a full scan. Partitioned and resumable index operations have sampling exceptions. Statistics refresh can explain a rebuild’s benefit, but page density and other factors can also matter.

Rebuilding that index does not refresh separate column statistics. In my test, the rebuilt index came back fully sampled. An automatically created statistic on another column still showed 353,333 modifications since its last update.

If your weekend job rebuilds indexes for that side effect, look at the clock. A statistics job that runs in minutes can buy the same thing. I wrote about the rest of that habit in Your Index Rebuild Maintenance Plan Is Rebuilding Indexes Nobody Uses.

Why This One Is Worth Your Attention

It shows up as randomness, and randomness is the hardest thing to get a team to investigate. The query is fast in the morning and slow in the afternoon. It’s fast in test and slow in production with identical code. It was fine for a year, and then it wasn’t.

People start blaming the network, the storage, the other team, the phase of the moon. Nobody suspects the small summary object that nobody has looked at since the database was built. Update the statistic, then check the estimate and the reads again.

The optimizer is not guessing badly, it is answering from a description of your data nobody has corrected.

Published by Pinal Dave on SQLAuthority. More of my work at pinaldave.com.

First appeared on Outdated Statistics in SQL Server: The Optimizer Is Planning for Last Year

Read the whole story
alvinashcraft
8 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Gothic Fiction: 8 Essential Elements Every Writer Should Know

1 Share

Discover the elements of Gothic fiction that create a dark, unsettling atmosphere and keep readers intrigued from beginning to end.

The term ‘Gothic fiction’ was first used in 1764, when Horace Walpole published The Castle Of Otranto. Since then, the genre has grown to include stories by writers such as Mary Shelley, Shirley Jackson, Stephen King, and Eudora Welty.

Gothic fiction explores fear and the darker side of human experience. It often features unsettling places, troubled characters, family secrets, and the consequences of the past. A character may be haunted by guilt, grief, memory, or regret rather than by a ghost.

What Is Gothic Fiction?

Gothic fiction creates an atmosphere of mystery and unease, often using isolated or unfamiliar settings to make characters feel vulnerable. Stories may take place in abandoned houses, crumbling mansions, castles, churches, graveyards, ruined estates, or isolated towns.

Supernatural events are common, but they are not essential. A story can create a sense of the Gothic through suggestion, psychological conflict, strange events, or a place with a disturbing history.

8 Essential Elements Of Gothic Fiction Every Writer Should Know

1. Gothic Architecture

Gothic settings create atmosphere and can make characters feel isolated or vulnerable. Castles, mansions, churches, graveyards, ruins, and abandoned buildings are common, but even an ordinary room can become Gothic if it feels threatening or has a disturbing history.

Shirley Jackson’s The Haunting Of Hill House and Bram Stoker’s Dracula use large settings to create unease. Stephen King’s 1408 shows that you do not need an entire mansion or castle—a single hotel room can become a Gothic space.

2. The Past & The Present

Gothic fiction often shows how the past continues to affect the present. Guilt, grief, secrets, scandals, and painful memories can shape characters and create conflict.

In The Picture Of Dorian Gray by Oscar Wilde, Dorian’s portrait becomes a reminder of the consequences of his actions. In Stephen King’s Needful Things, characters’ desires expose secrets and lead to terrible consequences.

The past rarely stays buried in Gothic fiction.

3. Eerie Elements

Eerie details create a feeling that something is wrong, even when the reader cannot explain why. Strange sounds, unsettling images, unusual behaviour, and unexplained events can build dread without relying on shock or violence.

Frankenstein by Mary Shelley and Oliver Twist by Charles Dickens show that Gothic unease does not depend on traditional ghosts. Dark settings, disturbing events, and emotional intensity can be enough to make readers feel that something is not right.

4. Tension & Anticipation

Tension keeps readers turning the pages because they want to know what will happen next. Gothic fiction builds this tension through uncertainty, delaying answers and dropping small clues—a locked door, footsteps in an empty hallway, or a character reacting strangely to an old photograph.

These details create unease before the reader understands the danger. Robert Bloch’s Psycho, later adapted into a famous film, draws on Gothic traditions to build fear through suspense, isolation, and uncertainty.

5. Gothic Creatures

Gothic creatures give writers a way to explore fears about death, desire, identity, and the unknown. Vampires, werewolves, mummies, ghosts, and monsters are familiar examples, but Gothic stories can also use strange or frightening human figures.

Bram Stoker’s Dracula uses the vampire to explore fear and desire, while Robert Louis Stevenson’s The Strange Case Of Dr Jekyll & Mr Hyde uses transformation to explore the darker side of human nature.

Vampires, werewolves, mummies, ghosts, and monsters have all found a home in Gothic fiction.

6. Decay & Decline

Decay reminds readers that nothing lasts forever. Crumbling buildings, fading beauty, broken families, and lost power can suggest death, secrets, and the consequences of the past.

Stephen King’s The Shining uses the isolated Overlook Hotel to create a sense of decay and menace. In The Picture Of Dorian Gray, the portrait becomes a powerful image of moral and emotional decay.

7. Time & Culture

Gothic writers can use the culture and details of a particular time to make the familiar feel strange or unsettling. Music, art, fashion, architecture, technology, customs, and social attitudes can all create atmosphere and a strong sense of place.

If your story is set in the past, research the period carefully. If it is set in the present, look closely at the world around you. Familiar details can evoke nostalgia, loss, or unease when they remind us that times change.

Stephen King’s It uses familiar elements of small-town American life in the 1950s and 1980s, from schools and shops to television and childhood games, to make the ordinary feel disturbing. The recognisable setting makes the horror more immediate.

8. Gothic Sub-Genres

Knowing the different Gothic sub-genres can help you decide what kind of story you want to write. Southern Gothic, Gothic Romance, Contemporary Gothic, and Gothic Horror all use Gothic elements in different ways.

Southern Gothic draws on the history and culture of the American South. Gothic Romance combines dark settings with romantic and emotional tension. Contemporary Gothic brings Gothic themes into the modern world, while Gothic Horror focuses more strongly on fear, supernatural threats, monsters, and psychological terror.

Examples

Southern Gothic: To Kill a Mockingbird by Harper Lee
Gothic Romance: Jane Eyre by Charlotte Brontë
Contemporary Gothic: Mexican Gothic by Silvia Moreno-Garcia
Gothic Horror: The Shining by Stephen King

Understanding these variations can help you find the Gothic style that suits your story.

The Last Word

Gothic fiction gives writers many ways to create mystery, atmosphere, and suspense. You can build a story around a place, a character’s past, a supernatural threat, or something that cannot quite be explained. We hope these ideas inspire you to explore the darker side of fiction and write your own Gothic story. You may also enjoy reading The Essential Guide To Writing Horror Stories.

Photo by Peter Herrmann on Unsplash

By Alex J. Coyne. Alex is a writer, proofreader, and regular card player. His features about cards, bridge, and card playing have appeared in Great Bridge Links, Gifts for Card Players, Bridge Canada Magazine, and Caribbean Compass.

If you enjoyed this, read these:

  1. What Is Genre? A Guide To The 17 Most Popular Fiction Genres
  2. Here Be Dragons – 7 Tips For Writing Dragons In Fiction
  3. What Is Dystopian Fiction & How Do I Write It?
  4. The 5 Pillars Of Speculative Fiction
  5. The Strange World of Slipstream Fiction – What Is It?
  6. How To Write Alternate History: A Complete Guide For Writers
  7. The 4 Pillars Of Magic Realism
  8. What Is Steampunk? How To Write Steampunk Fiction

Top Tip: Sign up for our free daily writing links.

The post Gothic Fiction: 8 Essential Elements Every Writer Should Know appeared first on Writers Write.

Read the whole story
alvinashcraft
8 hours ago
reply
Pennsylvania, USA
Share this story
Delete

The Test Was Wrong. Rewriting.

1 Share

Did you ever see this line when you let your code agent run wild?

“The test was wrong. Rewriting.”

Can the genie really make mistakes?
Ok, enough jokes.

If you’re surprised, this may be the first time you’re generating tests. It happens a lot – seven times in a single build in one of my checks.

Let’s walk through a couple of scenarios, and see how we got to this place.

For our purpose, let’s assume that we generated both code and tests. But the analysis is true for when the code agent generates just the test.

How did we get here?

The genie generates test code. This code is just lines of text in a file – unproven until it runs. It can be proven wrong syntactically, and by some heuristics, logically. Meaning, knowing it’s wrong without running it.

So if the genie runs this check only (we don’t really know unless it decides to tell us), it’s a guess. Not 50-50, but still not a real proof.

Our genie is nothing but a truth chaser. Rightly so, it runs the test. If the generated test fails, we have a fork in the road.

  1. The code is right, and the test is wrong.
  2. The test is wrong, and the code is right.

Nah, it’s not that simple. Both can be wrong. It can also be:

  1. The genie understands the test doesn’t even run the code. Or,
  2. The genie realizes the test isn’t designed correctly to check what it needs to.

Happened to me, where a test that should have recreated a race condition turned out to never run the race at all.

I want you to understand that at least in one of the stations, the genie made a mistake.

  1. Creating the code
  2. Creating the test
  3. Validating the test vs the code
  4. Running the test
  5. Evaluating the test for its purpose

#4 is the easiest to spot. Why?

Because it’s not reasoning. It’s a deterministic run – pass or fail. So a failure means something is definitely wrong.

In all other stages, whatever the genie does with the code is
a) Hidden from us
b) Non-deterministic, meaning we got the result today, but maybe if we ran this yesterday we’d get a different answer, and
c) It actually caught something. Or thinks it did.

Everything except the actual run is opaque and not repeatable. Not what we’d call proof.
A crime may have happened here. Or not.


This is not an anti-genie rant.

Let me be honest. I don’t – I can’t – read everything it logs. I run the agent, and look at the end results. Tests passing? Cool. I usually don’t look back at the logs.

Heck, the race condition issue? I asked Claude to go through the build logs to find where tests were found wrong.

I use coding agents all the time. And looking at me, you’d say – I trust them. Because we were taught that trust looks exactly like this.

But, this is not trust. It’s a bet. Lots of them. The app does work mostly, and when I find something I ask for a fix. And if the agent finds something, it fixes it.

But remember the options? The genie can make mistakes. Also in fixes. And in new code. And in replacement tests.

Who says the fix is the correct one?

My app, even in production, does not carry the risks of fully scaled apps. Imagine hundreds of developers, each with their own agent making those bets on your finance apps. Or law practices. Or online election management.

Are you scared? I know I am.

The way out is chunking the tasks to be smaller and manageable. Smaller pieces of code produce smaller logs. Ones we can read better.

Reviewability – if it’s not a word, it should be – is now a delivery capability.


Building that capability into a team – smaller pieces, logs someone actually reads, verdicts that get checked – is hard to do on your own. That’s where I come in – Mentoring and training on AI quality.

Built on your codebase and your quality goals. Mentoring for the ongoing work with your team, training when the ground needs covering.
Tell me what you’re dealing with. A couple of lines is enough.

The post The Test Was Wrong. Rewriting. first appeared on TestinGil.
Read the whole story
alvinashcraft
8 hours ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories