How to stop AI coding agents from re-reading the same repository files
AI coding agents stop repeating repository searches when you preserve verified discoveries outside the session, then retrieve them before the next investigation.
Do not save raw searches or whole transcripts. Save the conclusions a future agent would otherwise have to reconstruct: architecture, dependencies, ownership, constraints, and the reasoning behind unusual code.
Falconer is a context engineering platform that writes and updates your docs as your code changes, and gives coding agents one place to retrieve and update shared knowledge through the Model Context Protocol. The next agent can retrieve what the last one verified and saved.
TL;DR
- Agents repeat repository exploration when useful conclusions remain trapped inside one session or tool.
- Exploration is most of the run, not a preamble to it. Measured across 730 SWE-Bench Pro trajectories per model, GPT-5 spent 61% of its pre-edit steps on information gathering and Sonnet 4.5 spent 49%, with the first edit arriving halfway through the run for GPT-5 and a third of the way through for Sonnet 4.5.
- Save verified architecture, dependencies, decisions, ownership, constraints, and failure modes, not raw searches or transcripts.
- Retrieve existing knowledge first, then search the repository only for what is still unknown.
- Current code remains the source of truth for implemented behavior.
- Prompt caching reduces the bill for a repeated read. It does not stop the read from consuming the context window.
- Measure a repeat-read rate before and after, or you cannot tell whether the retrieval layer is helping.
- Falconer is a context engineering platform for people and coding agents that compounds across tools, repositories, and sessions.

Why do AI agents repeat repository searches?
Most coding tasks open with the same questions:
- Where is the entry point?
- Which service owns this behavior?
- What writes to this table?
- Where are the relevant tests?
- How does this code reach production?
- Why does this unusual constraint exist?
Exploring the repository is necessary the first time. The waste starts when every new session repeats it.
How much of a run is exploration rather than work has been measured. Nilenso’s analysis of agent trajectory shapes, covering 730 SWE-Bench Pro trajectories per model, found GPT-5 spending 61% of its pre-edit steps on information gathering and Sonnet 4.5 spending 49%, with mean trajectory lengths of 59.5 and 77.5 steps. In these benchmark trajectories, the first edit arrived 50% of the way through the run for GPT-5 and 35% for Sonnet 4.5.
Anthropic’s own Claude Code best practices names the pathological version of this directly, as “the infinite exploration”: you ask Claude to investigate something without scoping it, and it reads hundreds of files, filling the context. The same page notes that “a single debugging session or codebase exploration might generate and consume tens of thousands of tokens.”
Discovery, interpretation, and validation
Agents repeat three kinds of work:
- Discovery: Finding the relevant files.
- Interpretation: Understanding how those files work together.
- Validation: Checking that interpretation against documentation, tickets, discussions, and history.
Finding a file is usually the easy part. Reconstructing why the current design exists is what costs you.
Current code can show that a fallback exists. It may not explain that the fallback must remain until a migration finishes. That context may live in git history, architecture decisions, incident records, and project tickets.
We put numbers on this problem in why your AI agents burn tokens hunting for answers they should already have. In one internal Falconer example from agent loops are fine, retrieval is what you’re paying for, estimated input-token cost for a 20-turn conversation fell from about $9.50 to $3.20 after obsolete tool results were cleared and earlier turns were summarized.
The opposite failure is worse
Re-reading too much is the visible problem. Reading nothing at all is the quiet one.
A preliminary benchmark from Meetless on stale context tested whether agents check the repository before answering. In 26 of 40 cases, Haiku 4.5 returned the outdated answer having made zero tool calls. On the floor test, every one of ten models across three vendors wrote the fully stale version without reading a single file.
That is the boundary condition for everything below. A persistence layer that an agent trusts instead of checking is not an improvement, it is a faster way to be wrong. Saved knowledge narrows the search. It does not replace it.
What should you save after an investigation?
Save conclusions that stay useful after the current task closes.
Good examples:
- These files implement authentication.
- This worker is the only writer to the billing table.
- This configuration exists because of a production incident.
- Do not remove this fallback until the migration finishes.
- This service belongs to the Payments Platform team.
- This runbook defines the canonical deployment process.
- This approach was rejected because it breaks regional data isolation.
Five fields for every finding
Each conclusion should carry five things:
- The finding: What the agent learned.
- The evidence: Code, commits, pull requests, tickets, or discussions supporting it.
- The scope: Where the conclusion applies.
- The owner: Who should verify future changes.
- The review trigger: What would make the conclusion stale.
That turns repository exploration into organizational knowledge instead of temporary session context.
What should you not save?
Do not preserve everything an agent touches.
Leave out:
- Raw search results
- Every file the agent opened
- Temporary hypotheses
- Full reasoning traces
- Unverified conclusions
- File paths without explanation
- Facts already obvious from current code
- Task-specific details with no future value
Retrieval noise lengthens agent loops
A larger knowledge base is not automatically a better one. Low-quality context creates retrieval noise and longer agent loops.
Anthropic’s guide to effective context engineering for AI agents puts the target as finding “the smallest set of high-signal tokens that maximize the likelihood of your desired outcome.”
Clean retrieval beats feeding an agent everything you have.
Does prompt caching already solve this?
Partly, and only the part you can see on the invoice.
Prompt caching reduces the cost and latency of repeated input. It does not remove that input from the context window, so unnecessary context can still affect performance. A token spent re-reading a file you already understood is unavailable for reasoning about the change you are making, whether or not it was discounted.
Treat caching as a cost control and the retrieval loop as a capacity control. They solve different problems, and only one of them makes the agent smarter.
What repository context should persist?
Durable repository context usually falls into four categories.
Repository maps
A repository map explains the major packages, services, entry points, and runtime relationships.
It should answer:
- Which packages are deployable?
- Which packages are shared libraries?
- Where do requests enter the system?
- Where does persistent data live?
- Which services communicate?
- Where are tests and deployment definitions?
Keep it short. You are aiming for orientation, not a description of every directory.
Change-path guides
A change-path guide explains how to complete a common task safely.
Examples:
- Add an API endpoint
- Change a shared type
- Add a database migration
- Update a background worker
- Modify authentication
- Test a cross-package change
- Deploy a service
Each guide should name the relevant files, required checks, common failure modes, and reviewer or owner.
Decision records
Code shows the implementation. Decision records explain the tradeoffs behind it.
The format was popularized by Michael Nygard’s 2011 proposal for lightweight architecture decision records. The ADR community maintains additional guidance and templates.
Capture:
- The problem
- Constraints
- Options considered
- The chosen approach
- Why the alternatives were rejected
- Consequences
- Conditions that should trigger reconsideration
This matters most for code that looks unnecessarily complex without its history.
Verified findings
Some discoveries are smaller than a full architecture document and still worth keeping:
- One service is the sole writer to a table.
- A retry path excludes hard declines.
- A test fixture is shared across several packages.
- A compatibility layer remains because one customer has not migrated.
Store these as short, cited findings. Update or supersede them when the underlying system changes.
How do you build the retrieval loop?

Retrieve before you explore
Ask for the architecture, ownership, known constraints, related incidents, and canonical runbooks first.
The agent should retrieve what the company already knows before generating another explanation of it.
Search only for what is unknown
Use repository search, code reading, and git history to answer the remainder.
Falconer can inspect current code and use blame, commits, and diffs to trace authorship, implementation changes, and recorded evidence of intent. The capability is explained in the Falconer agent speaking git.
Validate before you save
Check the finding against current code and any relevant pull requests, tickets, documentation, or discussions.
A plausible interpretation is not durable knowledge until the evidence supports it.
Route high-risk findings to a human
Require review for anything affecting:
- Security
- Data integrity
- Compliance
- Production operations
- Major architectural boundaries
Lower-risk repository maps and development notes can use lighter review.
Write the conclusion back
Save the finding, evidence, scope, owner, and review trigger.
Falconer MCP supports targeted document updates, so an agent can change the relevant section instead of replacing the whole document.
Retrieve it next time
Future agents should read the relevant architecture and decision context before opening repository files.
Saved knowledge should narrow the search, not stop the agent from checking current code.
Re-check it when code changes
When implementation changes, identify which repository maps, guides, runbooks, or findings may now be wrong.
For eligible documents with auto-update enabled, Falconer analyzes merged pull requests and prepares targeted edits. Depending on the document’s update mode, Falconer either applies the edits or holds them for review.
How do you check that the retrieval loop is working?
Establish your own baseline before changing the retrieval loop. Pick one recurring task, run it the same way five times, and record four numbers.
- Total reads. Every file the agent opened.
- Distinct reads. Unique paths among them.
- Repeat-read rate. One minus distinct over total. This is the number the retrieval loop is supposed to move.
- Reads before the first edit. How much of the run happened before any work did.
Then check what actually loaded.
In Claude Code, /context lists the memory files in the session, which is how you catch a findings document that exists but never reached the model.
Run the same task again with retrieval in place and compare the four numbers on the same task, not on a different one.
Two cross-checks worth adding. Ask the agent to cite the source for each claim it makes about architecture, because an answer with no citation means it answered from priors rather than from your knowledge base. And plant a fact in the shared layer that is not derivable from the code, such as an ownership change, then ask a question that requires it. If the answer comes back without that fact, the loop is not connected.
Do not publish an improvement figure from a single task or a single week.
What does the agent prompt look like?
Give the agent the order of operations explicitly:
Before searching the repository:
1. Retrieve the current architecture for this subsystem.2. Find its owner, dependencies, known constraints, and related incidents.3. Identify the canonical runbooks and decision records.4. Search the repository only for unanswered questions.5. Validate new conclusions against current code and history.6. Propose updates for any reusable findings or outdated documentation.The fourth instruction matters most. Everything above it exists so the agent spends its search budget on questions nobody has answered yet.
Anthropic’s Claude Code best practices similarly recommend providing specific context, pointing the agent toward relevant sources, and separating exploration from implementation.
Why isn’t the retrieval layer helping?
When the loop is in place and the agent still greps its way through the repository, the cause is usually one of six things.
| Symptom | Likely cause | How to check |
|---|---|---|
| Agent never calls the retrieval tool | Retrieval is not in the instructions, only in your intent | Read the prompt back. Is step one “retrieve” or is it implied? |
| Agent retrieves, then searches anyway | The returned context did not answer the question it had | Log the query and the returned document. Was the answer actually in there? |
| Notes exist but never load | Wrong file location or wrong scope | /context in Claude Code lists what loaded |
| Agent contradicts the saved finding | The finding conflicts with current code and the agent is right | Re-validate the finding. Code wins on implemented behavior |
| Agent ignores a correct finding | The finding is undated, unsourced, or too long to trust | Add the evidence link and the review trigger |
| Retrieval returns too much | The knowledge base holds transcripts and raw searches | Apply the “what should you not save” list above |
The fourth row matters most. An agent that overrides a stale saved finding by checking the code is behaving correctly, and the fix is in the knowledge base, not the agent.
Is a bigger context window enough?
No. A larger context window can hold more files. It does not decide which files matter, explain why the design exists, or remove stale information.
Liu et al. found in Lost in the Middle that model performance can degrade when relevant information appears in the middle of a long context.
Chroma’s context rot study tested 18 models across four vendors and found substantially better performance with focused prompts of roughly 300 tokens than with full prompts of roughly 113,000 tokens. The Claude family showed the most pronounced gap. Filling the window is not the same as improving the answer.
Compare five context approaches
| Approach | Strength | Failure mode |
|---|---|---|
| Larger context window | Reads more code at once | Adds irrelevant or stale context |
| Repository instructions | Simple and local | Can become broad, duplicated, or outdated |
| Session or project memory | Preserves previous learning | Often stays inside one tool or repository |
| Shared retrieval layer | Returns relevant context on demand | Depends on maintained source knowledge |
| Shared read-write layer | Lets verified knowledge compound | Requires permissions and review controls |
The best setup gives the agent the smallest complete set of current, relevant context.
That is a retrieval problem before it is a model problem, which is the argument behind context engineering versus prompt engineering.
How do you measure less repetition?
Track:
- Repository searches per task
- Files opened before the first edit
- Time before the first meaningful change
- Repeated questions across sessions
- Existing context retrieved before exploration
- New context written back after investigation
- Corrections caused by stale instructions
- Documentation updates triggered by code changes
Treat improvements as internal benchmarks until you have enough data to publish them.
Establish a representative baseline
Choose two metrics and measure them across a representative set of recurring tasks.
Compare similar tasks before and after adding the retrieval loop. Do not assume one task or one week is enough to establish a reliable trend.
For more background on persistence models, how others build agent memory covers the designs evaluated while building Falconer’s agent memory.
How does Falconer fit?
Falconer is a context engineering platform that writes and updates your docs as your code changes. It connects code, documents, tasks, and conversations into a shared context layer that people and coding agents can retrieve.
With Falconer, teams can:
- Let Claude Code, Cursor, Codex, and other MCP clients search and read shared documents.
- Ask codebase questions and get cited answers tied to supporting sources.
- Read git history when the current repository snapshot does not explain a change.
- Write verified findings back as targeted document updates.
- Keep eligible architecture guides and runbooks connected to merged code changes.
- Reuse the same knowledge across Slack, the Falconer editor, coding agents, and the CLI.
Agents should still read code. What they should not do is rebuild the same understanding from zero every time.
Close the organizational loop
Falconer turns each investigation into an advantage for the next one.
Coding tools can preserve project instructions or local memory, but that context often stays inside one tool or repository. Falconer closes the organizational loop: agents can retrieve shared company knowledge, validate it against current code, and write durable conclusions back to the same source.
What compounds is the knowledge base itself, across people, tools, repositories, and sessions.
FAQ
Why does my coding agent repeat searches?
Each session begins with limited access to what previous sessions concluded. Some coding agents provide project instructions, resumable sessions, or local memory, but those mechanisms do not automatically create shared organizational knowledge across people, tools, repositories, tickets, and documents.
A shared retrieval layer makes verified conclusions available outside the session that produced them.
Does a larger context window solve this?
No. A bigger window lets an agent read more code in one pass. It still has to determine which code matters, identify stale information, and recover context that lives outside the repository.
Retrieval that returns a small, relevant set of current evidence is often more useful than loading the whole repository.
Can repository instructions solve this?
Yes, for repository-scoped context.
Files such as CLAUDE.md, .claude/rules/, Cursor project rules, and AGENTS.md can preserve commands, conventions, architecture decisions, workflows, and other project knowledge.
Claude Code’s memory documentation also describes project-scoped auto memory, while Cursor rules provide reusable context for Cursor agents.
These tools do not create one shared organizational memory across coding tools, repositories, documents, tickets, and conversations. Falconer is useful when context must be shared across people and agents, supported by sources, and maintained as the organization changes.
Should saved knowledge override current code?
No. Current code is the source of truth for implemented behavior.
Saved knowledge explains architecture, recorded intent, ownership, and constraints. When the two conflict, the agent should flag the contradiction and inspect the evidence rather than silently choosing one.
How do saved findings stay current?
Attach a review trigger to each finding:
- Service changes
- Schema changes
- Configuration changes
- Dependency changes
- Deployment changes
- Ownership changes
For eligible documents with auto-update enabled, merged pull requests can trigger targeted edits. Depending on the update mode, Falconer either applies them or holds them for review.
Should every investigation become documentation?
No. Preserve verified findings with likely future value and let the rest go.
A one-off debugging hypothesis should disappear with the session. A confirmed failure mode, architectural constraint, ownership discovery, or rejected approach should not.
How do you compare retrieval with grep?
Run both against the same real tasks.
Track:
- Files opened before the first correct edit
- Time before the first correct edit
- Repeated questions across sessions
- Corrections caused by missing context
- Evidence retrieved outside the repository
Grep is strong for exact repository lookups. It cannot answer questions about decisions, incidents, ownership, or product constraints unless that information exists in the repository.
What should a team capture first?
Start with whatever agents investigate repeatedly:
- Authentication and authorization
- Data ownership and write paths
- Deployment and rollback
- Shared package dependencies
- Common change paths
- Known production constraints
- Service ownership
- Rejected architectural approaches
A small set of current, high-value guides will reduce more repeated work than a large collection of unreviewed agent summaries.
Ready to get started?
Create an account and start building your knowledge base — no contracts or credit card required. Or, contact us to design a custom package for your team.

Ready to get started?
Create an account and start building your knowledge base — no contracts or credit card required. Or, contact us to design a custom package for your team.