Context Engineering: Why Giving Your AI More Information Can Make It Worse
There is an assumption built into most early AI architecture decisions: if the model has access to more information, it will produce better answers. Add more documents to the knowledge base, increase the context window, connect another data source, and retrieve more chunks. The logic feels sound. In practice, it breaks down in ways that are easy to miss until the system is already in production.
A Chroma Research study published in July 2025 tested 18 frontier models, including GPT-4.1, Claude 4, and Gemini 2.5, and found that every single one performed worse as input length increased, even on simple tasks. The researchers called the phenomenon "context rot": the steady decline in output quality as the context window fills, arriving well before the token limit does. More tokens in, worse answers out, across every model tested.
That finding matters for enterprise AI because the instinct when a system underperforms is usually to add more context. The research suggests the instinct is often wrong.
The Context Window Is Not Storage
Modern language models can process increasingly large amounts of text, which removes a technical constraint but can create a misleading assumption: if information fits inside the context window, there is little downside to including it.
Consider an internal AI assistant tasked with explaining a company's travel reimbursement policy. The retrieval system surfaces the current policy, two previous versions, several employee FAQs, an old email about an exception, meeting notes on a proposed change, and a document from another regional office. Every document is related to the question. None of them is obviously wrong to retrieve. And yet the model now must determine which source is authoritative, which version applies, whether the exception is still valid, and which regional rules are relevant. At that point, the bottleneck is not access to information. It is the model's ability to determine which of its actual applications applies.
This is the core challenge that context engineering addresses. The question is no longer how much information a model can process. It is what information should reach it at a particular moment, from which source, and under which conditions.
Similarity and Relevance Are Not the Same Thing
Retrieval systems typically rank information by semantic similarity, which is useful but incomplete. A document can use almost the same terminology as the user's question, while being outdated or applicable to a different process. Another document may contain the authoritative answer but use different wording and rank lower as a result.
When both are retrieved, the model must resolve the ambiguity rather than simply answer the question. Five highly relevant chunks will consistently outperform fifty loosely related ones. This is why production RAG systems increasingly require relevance thresholds, recency signals, metadata filters, and deduplication rather than simply returning more results. Anthropic's own Contextual Retrieval approach reduced failed retrievals by 49%, and by 67% when reranking was added on top.
The implication is significant: retrieval quality is a separate engineering problem from model quality, and improving one does not automatically improve the other.

Outdated Information Can Be More Damaging Than Missing Information
Enterprise knowledge changes continuously. Policies are revised, products are updated, procedures are replaced. The previous versions, however, often remain perfectly searchable.
If a retrieval system surfaces both version 4 and version 5 of an operational procedure, and version 4 matches the user's wording more closely, the model may produce a coherent and confident answer based on the wrong procedure. Nothing failed at the model level. The model reasoned correctly over the context it received. Replacing the model does not change what the retrieval layer sends it.
According to a 2026 study by Cloudera and Harvard Business Review, only 7% of enterprises say their data is completely ready for AI. Well-structured context built from stale definitions is worse than a messier context from fresh sources, because it looks authoritative. A model that confidently surfaces outdated information is harder to catch than one that admits uncertainty.
Context Is Also an Access Control Problem
An enterprise AI system may connect to CRM platforms, internal documentation, ticketing systems, file repositories, and employee records. Giving it broad access can make it more capable in some situations. But capability and appropriateness are different questions.
A customer support agent may need order history but not a customer's complete account record. An internal search tool may technically retrieve confidential documents that certain users are not authorized to access. Context engineering, therefore, intersects directly with identity, permissions, and data governance. The relevant question is not whether the system can retrieve a piece of information, but whether that information should be available to this system, for this user, for this task, at this moment.
Treating access control as an afterthought, something to be applied after retrieval rather than built into it, is one of the more common architectural mistakes in enterprise AI deployment.
What Teams Are Doing About It
The most effective approaches treat context quality as an engineering problem rather than a configuration problem. Several patterns appear consistently in production RAG systems:
Relevance thresholds and reranking. Rather than retrieving the top-k results by similarity score alone, teams add a second pass that evaluates whether each chunk is genuinely useful for the specific query. Chunks that are semantically close but factually stale or contextually mismatched are filtered out before they reach the model.
Recency and authority signals. Retrieval pipelines increasingly weight documents by how recently they were validated and which source takes precedence when information conflicts. A policy document updated last week should outrank one from two years ago, regardless of embedding similarity.
Access-aware retrieval. Permissions are applied at the retrieval layer rather than after the fact. The system does not retrieve documents that the current user cannot access; rather, it retrieves everything and filters the response.
Context budgeting. Teams define how much of the context window each component can use: system instructions, retrieved chunks, conversation history, and tool results. Without explicit budgets, conversation history tends to grow until it crowds out the information that matters most.
None of these requires a different model. They require a different architecture around the model.
Sufficient Context, Not Maximum Context
Nearly 65% of enterprise AI failures in 2025 traced back to context drift or memory loss, not model capability issues or weak training data. The agent simply lost track of what it was doing because its context window filled with information that competed with what mattered.
Every additional piece of information the model receives must be processed. Longer prompts increase token consumption, response latency, and inference costs. Removing too much context can leave the model without what it needs. Adding everything available can make the system slower, more expensive, and less reliable. The goal is sufficient context, not maximum context.
A well-designed context layer is selective by design: relevant, current, authoritative, and scoped to the task's requirements. The decisions made at this layer, about what to include, what to filter, which source takes precedence, and what the user is authorized to see, can have as much influence on system behavior as the model itself.
At ASSIST Software, context architecture is one of the areas where we consistently see the largest gap between AI prototypes and production systems. The retrieval layer, the permission model, and the information governance decisions made early in a project shape what the system can reliably do long after the initial deployment. If you are working on systems where information quality matters as much as model quality, we would like to hear about what you are building.

Frequently Asked Questions
What is context engineering in AI?
Context engineering is the discipline of deciding what information reaches an AI model, when, from which source, and under which conditions. It goes beyond prompt engineering to address the complete information environment in which the model operates: retrieved documents, conversation history, system instructions, permissions, and business rules. In enterprise AI systems, context engineering decisions can influence output quality as much as the model itself.
What is context rot in large language models?
Context rot is the measurable decline in an AI model's output quality as its input context grows longer, even when the task itself stays equally complex. The term was coined by Chroma Research in a July 2025 study that tested 18 frontier models and found every single one degraded with longer inputs. The implication for production systems is that adding more information to the context window does not reliably improve performance and can even worsen it.
Why does retrieval quality matter as much as model quality in RAG systems?
In retrieval-augmented generation systems, the model can only reason over the information it receives. If the retrieval layer surfaces outdated, duplicated, or weakly relevant content, even a highly capable model will produce answers based on the wrong context. Improving the model will not fix a retrieval architecture that cannot distinguish between relevant and authoritative information. Retrieval quality and model quality are separate engineering problems that both require deliberate attention.



