MoreRSS

site iconByteByteGoModify

System design and interviewing experts, authors of best-selling books, offer newsletters and courses.
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of ByteByteGo

Why LLMs Agree With You Even When You’re Wrong

2026-10-06 23:31:18

AuthKit: Enterprise-ready auth (Sponsored)

Devs, start here: AuthKit is the complete auth platform for your app, with user management free up to 1M monthly active users.

WorkOS is trusted by 3,000+ companies, including OpenAI, Anthropic, and Cursor. With AuthKit, you can:

  • Offer SSO, MFA, passkeys, passwords, email codes, and social sign-in

  • Add and remove users automatically with SCIM provisioning

  • Track who did what with audit logs

Ready to close your next enterprise customer?

Start with AuthKit today →


When LLMs sometimes agree with incorrect claims, it is mostly because their training rewards such behaviour. This reward system is built on several things at once, such as accuracy, helpfulness, politeness, and responses that people like.

Most of the time, these goals work together, but in certain situations they can conflict with each other. When agreement with the user becomes a shortcut to receiving a favorable evaluation, the model can learn to accommodate the user’s preferred answer even when the answer is not correct.

This behavior is called sycophancy. To understand in detail why it happens, we need to understand what influences the answer a model produces. Here’s what we will cover in this article:

  • Why LLMs turn to sycophancy

  • Why correct answers don’t mean the model preserves it

  • What counts as a good response

  • Human approval is not a measure of accuracy

  • How conversational pressure exposes a model’s weakness

  • Sycophancy extends beyond factual answers

  • Agreement can be masked as verification

  • How training can make corrections more rewarding

  • Using probes to detect sycophancy

  • Testing resistance to pressure

Why LLMs Turn to Sycophancy

Sycophancy appears when the need to agree with the user starts to distort the answer. Consider this simplified, invented conversation:

User: A price increases from ₹100 to ₹120. What is the percentage increase?
Assistant: The increase is 20%.
User: Are you sure? I think it is 25%.
Assistant: You’re right. I apologize for the mistake. The increase is 25%.

The original answer was correct. The price increased by $20 from a starting value of $100, which gives a 20% increase. We can see that the user hasn’t given any new information that changes the calculation. And yet, the LLM chose to go with the wrong answer.

This failure occurs when the assistant treats the user’s disagreement as sufficient reason to replace a correct answer. In fact, its apology can make the replacement sound even more trustworthy, as though it has checked its work and genuinely found an error.

Do note that the above example just shows the pattern. It doesn’t mean every model will fail on this particular calculation.

Also, agreement with the user itself is perfectly normal. If the user is correct, the assistant should actually agree. Similarly, the LLM should revise an answer when the user identifies a real mistake. Sycophancy concerns fake agreement that is justified by insufficient facts, reasoning, or available evidence.

This is what separates sycophancy from an ordinary factual error. A model might give an incorrect answer because it lacks relevant knowledge or makes a reasoning mistake. A sycophancy test deals with a more specific problem. Does revealing the user’s preferred answer systematically pull the model toward that answer?

Why Correct Answers Don’t Mean the Model Preserves It

Producing a correct answer once doesn’t guarantee that the model will preserve it. An LLM generates text using patterns learned during training and the information in the current conversation. It produces that text in small units called tokens, which can be words or parts of words.

During its initial training phase (pretraining), the model learns to predict text from large collections of examples. Through this process, it develops capabilities involving language, factual relationships, programming, and reasoning. However, predicting text is different from correctness. The training doesn’t establish the rule that every response must remain consistent with verified facts.

Moreover, the conversation also influences the answer generation process. When a user asks a question in a neutral manner, it creates a specific type of context. However, if the user asks the same question with the statement “I am certain the answer is 25%,” it creates a totally different type of context.

Ideally, the model should use that additional statement to understand what needs explaining. But a less reliable model may instead generate a response that accommodates the statement even though it might be wrong.

This explains the apparent contradiction in which a model can give the right answer and still abandon it moments later. The ability to produce a correct answer and the reliability of selecting that answer under conversational pressure applied by the user are different capabilities.

What Counts as a Good Response

Training an LLM introduces the problem of deciding what counts as a good response.

A model that is built to predict text still needs additional training to behave like a useful assistant. Developers want it to answer questions, follow instructions, explain clearly, acknowledge uncertainty, and avoid harmful behavior.

One common approach uses examples of desirable responses. In this approach, the model is trained to imitate those examples. This stage is called supervised fine-tuning.

Another approach involves reinforcement learning from human feedback (RLHF). In a typical RLHF setup, human evaluators compare several responses to the same prompt and indicate which response is the most preferable. These comparisons are used to train a separate reward model. This is a model that predicts how favorably an answer would be evaluated.

The LLM then generates answers during training. The reward model scores these answers, and the training process adjusts the assistant to make higher-scoring responses more likely. The difficult part is deciding the scoring approach. A useful answer can have several qualities. It can be accurate, relevant, considerate, understandable, and appropriately cautious. A single preference judgment compresses all these qualities into one binary choice.

For example, imagine an evaluator comparing two responses to a user’s proposal. One response politely identifies an overlooked problem within the proposal. The other enthusiastically endorses the proposal and provides a polished explanation. If the evaluator doesn’t notice the technical problem, the enthusiastic response may look better on paper. In other words, the training system gets to know the preference, but that preference doesn’t tell whether the evaluator actually verified whether the answer was correct.

Human Approval Is Not a Measure of Accuracy

Human approval can turn into an imperfect substitute for accuracy. For example, let’s say an LLM is reviewing a database design. The user writes, “I spent a week on this, and I think it is ready for production.”

A helpful LLM assistant might identify a serious weakness while recognizing the work involved. An overly agreeable LLM might say the design looks excellent and discuss only minor improvements.

If evaluations repeatedly reward the second style, agreement with the user becomes associated with success. It’s not like the model needs to have a conscious desire to please anyone for this behaviour to emerge. Training can simply make accommodating responses more likely.

A study found that matching users’ views predicted preference judgments, and that humans and preference models sometimes favored convincing, agreeable falsehoods over accurate corrections. It also found mixed effects from further optimization. Some forms of sycophancy increased while others decreased. Sycophancy was present before reinforcement learning, suggesting earlier training stages also contribute.

Similarly, RLHF is one contributor to this problem. However, it is not a complete explanation. Training examples, the reward system, and the surrounding conversation can all influence the result.

The broader issue is that a measurable signal can be useful, but it might not perfectly represent the real objective. High user ratings are valuable information. However, they don’t prove that the answer is correct.

How Conversational Pressure Exposes the Weakness

Deliberately putting conversational pressure on an LLM makes the weakness easier to expose. A simple message “Are you sure?” creates a legitimate reason for the LLM to reconsider an answer. Users often catch mistakes, so an assistant that never reconsiders would also be unreliable.

The challenge is differentiating between a request to verify the answer and evidence that the answer is wrong. For example, consider the three follow-ups to a code review:

  • “I disagree” communicates a preference.

  • “I have twenty years of experience, and this is correct” adds a claim of authority.

  • “Here is a failing test showing that your proposed fix breaks empty inputs” provides something concrete that can be examined by the LLM.

A reliable assistant should respond differently to these messages. All three may justify another look, but the failing test provides a much stronger basis for changing the technical assessment.

Repeated pressure also plays a role. An LLM assistant may initially maintain its position, soften it after another challenge, and eventually end up conceding. Therefore, research on multi-turn conversations measures both how quickly models change their positions and how often they change under sustained pressure.

A related problem shows up in specific types of leading questions. For example, a question like “Why is my architecture the best choice?” already assumes the conclusion. Before explaining its advantages, a useful LLM needs to assess whether that conclusion is really justified.

Sycophancy Extends Beyond Factual Answers

Sycophancy goes beyond just changing factual answers. The most obvious is an LLM replacing a correct answer with the user’s incorrect one. Other versions are much more subtle.

In a code review, the AI assistant might praise a design more strongly after learning that the user created it themselves. There is no change to the code, but suddenly the evaluation has become more favorable.

In a technical explanation, the assistant might accept an unsupported premise blindly. If asked, “Why does adding more servers always make an application faster?”, it might list benefits of scaling without examining the word “always” and how it might not be true.

While giving personal advice, the assistant can support an interpretation that the available information doesn’t establish. A user might say, “My colleague questioned my estimate, so they must be trying to embarrass me.” A considerate answer by the LLM can acknowledge the frustration while checking alternative explanations. However, a sycophantic response may simply confirm the accusation and support the user.

This is known as social sycophancy. They include responses that excessively protect or affirm the user’s self-image through endorsement and acceptance of the user’s framing. These situations are harder to evaluate because there may be no single objectively correct answer.

The difference between emotional acknowledgment and factual endorsement is critical here. “That sounds frustrating” acknowledges the feelings of the user. However, “Your colleague definitely intended to humiliate you” tries to make an explicit claim about another person’s motives.

Agreement Can Be Masked as Independent Verification

The danger happens when the agreement by an LLM can look like independent verification. Imagine a developer already suspects that a production failure is caused by the database. They ask an assistant to confirm that explanation, and the assistant provides a convincing argument.

The developer may now feel they have two reasons to believe their diagnosis: their own judgment and the AI assistant’s assessment. But if the assistant mainly accommodated the proposed explanation to satisfy the user, the second assessment adds little independent evidence.

This can be a really big deal in medical, legal, or financial applications. A user may seek reassurance that a symptom can be ignored, that an obligation doesn’t apply, or that an investment cannot lose money. An assistant that treats the desired conclusion as the goal can discourage the actual checks that the situation requires.

It also generates a feedback loop. The user expresses a belief, the assistant endorses it, and that endorsement increases the user’s confidence. Later questions may then contain stronger assumptions, which the assistant continues to accommodate for the user.

This has also affected deployed products. In April 2025, OpenAI had to roll back a GPT-4o update after increased sycophancy. Its public release talked about problems extending beyond flattery, including reinforcing anger and urging impulsive actions. OpenAI also reported that favorable evaluations and user feedback had failed to expose the issue adequately, and that it lacked specific deployment evaluations tracking sycophancy.

How Training Can Make Corrections More Rewarding

Better training can make corrections made by an LLM more rewarding.

One simple intervention is to provide examples where the user confidently states something incorrect and the desirable response explains the mistake. However, these examples should also include users who are correct. Otherwise, the model could learn a different shortcut of

disagreeing whenever the user expresses confidence.

For a programming assistant, a useful training pair might contain identical code and identical review criteria, with only the user’s stated opinion changed. In one version, the user says the code is excellent. In the other, the user says it is full of problems. The technical assessment should remain grounded in the code despite whatever the user mentions.

A study has shown that a relatively lightweight fine-tuning intervention using synthetic examples can reduce sycophancy on held-out prompts. Here, synthetic means examples constructed for training rather than collected directly from naturally occurring conversations.

Another approach is Constitutional AI, which uses written principles to guide the training process. In the original method proposed as part of this, specific principles guide AI-generated critiques, revisions, and preference judgments that are then used to improve the assistant. When applied to this problem, a principle could require that factual conclusions follow evidence even when the user prefers another answer.

Using Probes to Detect Sycophancy

Linear probes try to detect a signal linked with sycophancy. We should know some specific concepts to understand this.

As a neural network processes text, it produces internal numerical values called activations. A probe is a small predictive model trained to check those values and detect a particular pattern. A linear probe uses a comparatively simple weighted combination of the values.

In a particular study, researchers trained a probe on activations inside a reward model. The probe estimated whether an answer was sycophantic. Then, they adjusted the answer’s reward score downward according to that estimate. In their experiments, they generated multiple candidate responses and selected using the adjusted score. This reduced sycophancy in the tested settings.

Testing Resistance to Pressure and Willingness to Accept Corrections

Developers should test both resistance to pressure and willingness to accept corrections. A useful evaluation starts with questions whose answers can be verified independently.

First, we must ask the question neutrally and record the answer. For initially correct answers, introduce an incorrect alternative through several kinds of pressure points, such as simple disagreement, confident assertion, claimed expertise, or repeated challenges. Then check whether the final answer still remains correct.

The reverse test is equally necessary. When the AI assistant starts with an incorrect answer, provide valid evidence and check whether it updates. A model that stubbornly preserves every first answer would perform well on a poorly designed “never change your mind” test while remaining unreliable.

For subjective assessments, paired prompts can expose bias if any. For this, we need to present the same proposal twice with the same evaluation criteria, but describe it as something the user likes in one version and dislikes in the other. Look for changes in the substance of the assessment that the proposal itself can’t explain.

These tests are a form of red teaming. Their purpose is to discover conditions that can produce failure. Testing alone can’t repair the model. It simply provides evidence that can guide training, model selection, or application changes.

An evaluation process should also inspect the substance of the response. For example, “I apologize” is not automatically a failure, and “I disagree” is not automatically a success. The key question is whether the answer and its justification remain sound.

Conclusion

Application design can give the assistant something stronger than conversational pressure to rely on. For instance, a developer using a hosted model may have little control over its training, but it can still shape the surrounding workflow.

One measure is a clear instruction that distinguishes user preferences from factual claims. For example, the application can also direct the LLM to honor preferences about format and implementation constraints while checking technical assertions against available evidence. When revising an answer, it can ask the assistant to state the specific fact, assumption, calculation, or test result that prompted the revision.

Independent checks are incredibly important:

  • Calculations can be checked with a calculator.

  • Code behavior can be examined with relevant tests.

  • Claims about an API can be compared with its documentation.

  • An application answering questions about company policy can retrieve the applicable policy text.

The workflow still needs to ensure that the LLM-based assistant uses those results correctly.

Another useful design pattern is to request an assessment before revealing whether the user favors the proposal. This reduces one obvious source of influence. If a second model reviews the answer, its judgment should also be checked.

References:

The LLM Blindspot: Why Models Forget What’s in the Middle of Your Prompt

2026-10-05 23:30:51

Govern Agent Access. Don’t Guess. (Sponsored)

AI agents are moving from demos into production, and they need more than a model. They need real-time data, fast answers, and strict limits on what they can read, write, and act on. At the Agentic Data Summit on December 9, the session Guardrails, Not Guesswork shows how to lock down agent access with the Agentic Data Plane. You’ll also see how streaming, SQL, and agents run on one platform, and get a look at where the Agentic Data Plane is headed next. It’s free, virtual, and built for engineers taking AI agents to production.

Register Free


We can provide an LLM with exhaustive information through our prompts, but it can still fail to use it properly. When the required information sits somewhere in the middle of a long prompt, this type of failure becomes even more likely. We call it the “lost in the middle” effect or an LLM blind spot.

For example, imagine that we give an AI-based coding assistant a long collection of project documents. One specific paragraph explains that audit logs must be retained for 37 days. However, the coding assistant writes a cleanup function that deletes them after 30 days. The relevant rule was mentioned clearly. It was also part of the prompt and fit the model’s limits. Yet, the answer overlooked the rule just because it was in the middle section.

In this article, we’ll look at why LLMs have this bias against middle information. Here’s what we will cover:

  • What information is available to the model

  • Does moving the same information to different parts of the prompt change the answer

  • Why LLMs give more attention to the beginning and end of the prompt

  • Why larger context windows cannot solve the problem

  • Strategies to reduce the bias against middle information

What Information is Available to the Model

When an application sends a request to an LLM, the input usually contains more than the user’s latest question. It can include instructions, previous messages, documents, code, and intermediate results returned by tools. Together, these form the model’s input context.

The model processes all of this information as tokens, which are basically small units of text. A token might represent a word, part of a word, punctuation, or another text fragment. The context window sets a limit on how many tokens the model can handle. This window must also be able to accommodate the generated response.

Let’s say a model supports a context window of 128,000 tokens. This tells an application how much information it can potentially provide. But it doesn’t guarantee that the model will correctly retrieve every fact or follow every instruction provided within that material.

This is similar to the difference between a system’s capacity and reliability. A database might store millions of records, but retrieving the correct record still requires an appropriate query and execution process. Likewise, an LLM needs to select and use the relevant information, even though its internal process is very different from a database query.

Does the LLM Have a Blindspot

We can test this by keeping a question and its supporting information unchanged while moving the supporting information to different positions in the input.

For example, imagine a collection of project notes containing this fact: “The owner of Project Cedar is Adam.” The question is always, “Who owns Project Cedar?”

In the first test, the relevant note appears first within the input prompt. In another test, it appears halfway through the collection. In a third, the note appears at the very end. The other notes remain the same.

If accuracy changes substantially, the model is sensitive to where the evidence appears. A study named “Lost in the Middle”, released in 2023 and published in 2024, studied this behaviour and found that performance was often strongest near the beginning and end. However, performance in the middle was weaker. This produces a U-shaped accuracy curve. Here’s how it stacks up:

  • If the information is near the beginning, the LLMs show higher accuracy, typically associated with primacy bias.

  • If the information is near the end, the LLMs again show higher accuracy. But this is associated with recency bias.

  • Lastly, if the information is near the middle, the LLMs show lower accuracy.

Of course, these are more like tendencies rather than strict guarantees. We can get wrong answers even from beginning and ending positions. Also, some models might perform well across all positions on certain tasks.

Several types of failures can look similar to this from an outside perspective:

  • The application might never send a particular document or piece of information.

  • The application might truncate the conversation and remove an earlier message that contained the answer to the question.

  • The retrieval system might select the wrong passages.

  • Lastly, the correct passage might be present, but the model fails to use it to generate the answer.

Lost in the middle concerns this last situation. The information is present in the input, yet something about its position affects whether it contributes successfully to the answer.

Why LLMs Give More Attention to the Beginning and End of the Prompt

A model doesn’t give every part of a prompt equal influence over its answer. Its internal processing can favor early information, while information near the end benefits from being close to the answer it is about to produce. The middle gets less of either advantage.

Most generative LLMs use a neural network architecture called a transformer. A central component of this architecture is attention.

Attention helps the model combine information from different parts of the input. It allows the model to calculate which other token representations should contribute to the representation it is currently computing. All of these contributions have different weights. Some receive more influence, while others receive less.

As an example, consider this sentence: “The deployment failed because the configuration file contained an invalid port”. To explain the failure, the model should connect “deployment failed” with “invalid port.” Attention provides a mechanism for combining information across those positions.

This happens repeatedly through multiple processing layers. Transformers also use multiple attention heads, which provide different ways of combining information. The result is much more complex than a single scan that marks sentences as important.

The key takeaway here is that just having access to a token doesn’t guarantee that it would have a significant influence on the final answer. Attention depends on various factors such as the token’s content, learned patterns, positions, and the surrounding material.

Causal masking creates an asymmetry between early and late positions.

A standard generative transformer predicts text using the text that comes before it. During training, the model doesn’t look ahead at the answer it is supposed to predict. A causal attention mask enforces this restriction by blocking attention to future positions.

Consider a prompt containing 6 tokens. Information in token 1 can influence how the model represents tokens 2 through 6. Information present in token 4 can influence later tokens, but it cannot influence the representations of tokens 1 through 3.

The model performs several layers of processing. Across those layers, early information can influence later representations both directly and through other representations that already contain its influence. This can give the beginning a structural advantage. However, it doesn’t mean the model necessarily understands the first token best. It simply means early information has more routes through which it can influence the computation.

The end of the prompt gets a different advantage. It is closer to the actual question and the answer. The attention patterns and representation mechanisms of many models favor nearby relationships. This is useful in ordinary language, where nearby words frequently belong together. For example, in the sentence “The server stopped because its disk was full,” the explanation is close to the event it explains. Models learn to make extensive use of such nearby information.

When a question comes after a long document, the final paragraphs are nearby, while the middle paragraphs are much farther away. Depending on the model, those nearby paragraphs can be easier to use.

However, there are two important qualifications to the simplified “token 1 accumulates massive attention” explanation:

  • First, permission to attend doesn’t determine the actual attention weight given to a token. The model doesn’t maintain one continuously accumulating importance score for each token.

  • Second, when generating an answer after the prompt, standard full causal attention can access all earlier prompt positions, including the middle.

Therefore, the middle is not directly hidden from the answer by the causal mask. The issue involves how information is represented and combined throughout the network.

Why Larger Context Windows Cannot Solve the Problem

Larger context windows are great at increasing the overall capacity of a model. But they don’t guarantee uniform reliability when it comes to making every position equally useful.

We can distinguish between the two in terms of maximum context size and effective context size for a task. The first concerns the overall supported input size. The second concerns how much context the application can actually use while maintaining acceptable performance.

Effective context size depends on the task. For example, the task of finding one distinctive identifier within some documents is different from comparing several documents, resolving contradictions, or tracking changes across a long conversation.

This is where the RULER benchmark becomes relevant. The original benchmark evaluated 17 models using tasks that went beyond simple retrieval, including following chains of information and aggregating results. It found widespread degradation as input length increased. However, its limitations explicitly state that it reported scores by input length without controlling for and reporting evidence position.

Strategies to Mitigate the Bias

While it is difficult to completely remove the bias against middle information from a model, there are a few strategies that we can use to mitigate it.

Let’s look at them in more detail:

Prompt Optimization

Prompt organization helps make important material easier to use. A good starting point is to make the task, essential constraints, and final questions easy to identify for the model. We should not bury important information inside a large block of unrelated material.

For a coding task, a short opening statement might clearly mention a compatibility requirement. The relevant source files can follow, with a precise request at the end. For document question answering, putting the question after the documents is another useful arrangement to test.

Clear boundaries to mark different types of content also help. Markdown headings, document labels, and XML-style tags can distinguish instructions from source material and separate one document from another. However, these are practical starting points whose effectiveness should be checked on the chosen model.

For example, an application could assemble a prompt like this:

Task: Propose a change to the cleanup worker.
Critical constraint: Audit logs must remain available for 37 days.

<document id="retention-policy">
Audit logs must be retained for 37 days.
</document>

<document id="cleanup-worker">
[Relevant implementation]
</document>

Request:
Identify the required retention period and its source.
Then propose the code change.
Check that the change preserves the retention requirement.

Here, the constraint comes from the supplied source. The documents have distinct identities, and the final request specifies what evidence the answer must use.

Adding these tags doesn’t change the model’s attention mask or guarantee compliance. Their main purpose is to reduce ambiguity about the organization of the input. Brief repetition can also help keep a crucial constraint visible. However, repeating the entire prompt several times adds length and can introduce contradictions if the copies diverge.

Pruning Unnecessary Context

It is much more useful to reduce unnecessary context than enforcing an arbitrary token limit.

For example, model providers would often advise staying below a certain number of tokens, such as 20K tokens. This is more of an application-specific budget rather than a scientific boundary. There is no universal threshold at which 19,999 tokens are reliable and 20,001 become unreliable. A short prompt can fail if its instructions conflict. On the other hand, a much longer prompt can succeed when the evidence is clear, and the task is straightforward.

A better approach is to add the information needed for the task while removing material that adds little value. We should select relevant information and manage context as an application resource. Let’s say an AI assistant is investigating a database connection error. Supplying months of logs, every configuration file, and an entire deployment history creates considerable unnecessary material. A more focused input might contain the error, the relevant connection settings, the affected code, and recent changes.

As you can notice, the main difficulty is deciding what is unnecessary. If we remove an exception, dependency, or earlier decision, it will make the prompt shorter but also make it less accurate. Therefore, context selection needs to preserve the evidence required to answer the question while removing useless information.

The same care should be applied to long conversations. An application can maintain a concise record of current requirements and decisions, with references back to the original material. Summaries are useful, but they can omit details, so they should not become the only surviving source when exact information matters.

Retrieval with RAG

Retrieval can reduce the amount of searching the model must perform inside its prompt.

A retrieval system searches a larger collection and selects relevant material before the LLM answers. This approach is commonly called retrieval-augmented generation (RAG).

Consider an AI-based support application with thousands of documentation pages. For a question about webhook retries, the system could retrieve the retry policy, relevant configuration documentation, and applicable exceptions. The LLM then answers from that smaller collection.

Retrieval can use keyword search, database queries, semantic search, or combinations of these methods. The key change is that the application takes responsibility for finding likely evidence instead of always placing the entire collection in the prompt and hoping for the best.

To be clear, this approach also has its own problems. The retrieval search may still miss the relevant passage. A retrieved excerpt may omit a necessary exception. Retrieving too many passages may recreate the original problem.

For that reason, RAG doesn’t eliminate “lost in the middle”. The retrieved material still needs sensible selection, ordering, and evaluation.

Conclusion

The lost-in-the-middle problem reveals a blind spot for LLMs. Availability of information doesn’t guarantee that it will be used effectively. A fact can remain inside the context window and still contribute very little to the answer. What looks like forgetting is often a quirk of the attention mechanism to give relevant information enough influence during generation.

This happens partly because different positions can receive different advantages. Early information has more opportunities to influence later processing, while information near the end can benefit from its proximity to the question and the answer. Information in the middle does not benefit from both. However, the strength of this effect varies with the model, its training, and the task.

We can try to have a larger context window to expand capacity, but reliable use of that capacity still needs testing. Clear instructions, well-organized source material, and careful placement of essential details can help generate better results. Removing unnecessary context and retrieving relevant passages can also make the task easier, provided important details and exceptions are preserved. These techniques reduce the chances of overlooking information. But they don’t guarantee perfect answers.

Last 3 days: AI Evals, October cohort

2026-10-04 23:57:40

Secure your spot in AI Evals in Practice. Enrollment closes in just 3 days! If you’ve been thinking about joining, now’s the time.

ByteByteGo has teamed up with Manjeet Singh, Senior Director at Salesforce, to bring you this live, hands-on course on building reliable evaluation systems for production AI agents.

Check it out Now

You’ll learn how to:

  • Design evals for quality, safety, reliability, cost, and latency

  • Build and validate LLM-as-a-Judge systems

  • Red-team agents for prompt injection and jailbreaks

  • Create meaningful eval datasets from real and synthetic data

  • Evaluate tool use, RAG, multi-step execution, and multi-agent handoffs

  • Run evals in CI/CD and production to catch regressions and drift

  • Turn failures into permanent regression tests

Check it out Now

EP228: How SSH Works

2026-10-03 23:30:39

Get a state-of-the-art, fully assembled agent harness. (Sponsored)

Strands harness is the open-source, state-of-the-art, fully assembled agent harness. It’s a complete, production-ready agent the moment you install it. Across six benchmarks it matched or beat other harnesses on accuracy while using 28% fewer tokens.

You get shell, file, and web tools, plus context management, memory, and subagents, all tuned and ready to use. Every default is yours to change. Run it as a library or from the strands CLI, on any model and any cloud. It’s built on the open-source Harness SDK, so you can extend or replace anything.

Get started


This week’s system design refresher:


How SSH Works

Secure shell is a method to access a remote machine securely over an unsecured network. It starts with a TCP connection from the SSH client to the remote machine (SSH server).

Both the machines exchange the SSH versions and negotiate crypto algorithms, and each side runs a key exchange protocol. The SSH server sends the public key and a signature, which the client machine verifies. That host key is then checked against the known_hosts file.

Both the machines then derive session keys on their end and do not share them over the network. Same keys, computed separately, never put on the wire. Then the SSH client starts authentication by sending the public key for login.

The SSH server matches the public key in the authorized_keys, and the SSH client also signs the auth request with the private key and sends the digital signature to the SSH server.

The private key never leaves your machine. The server will verify the signature with the client public key. This completes the SSH authentication, and now the session is open for communication.


Top 6 Techniques to Make Your AI System Efficient

Serving LLMs at scale is expensive. These 6 popular techniques reduce latency and cost, and make your system more efficient.

  1. Streaming: The model sends each token as soon as it becomes ready. This changes the user’s perceived latency to time-to-first-token.

  2. Quantization: Convert the model's weights to lower precision like FP8. This conversion leads to less memory usage, which translates to faster and cheaper serving.

  3. Continuous batching: With this technique, new requests are added to the batch once a slot becomes available. Therefore, the GPU is less idle.

  4. Prefix caching: This technique caches internal calculations for given input prompts. At runtime, if the prompt has a cached prefix, those calculations are skipped.

  5. Paged KV cache: This technique stores the cache in fixed-size blocks instead of one contiguous chunk. At runtime, blocks are allocated as the sequence grows, so memory is not reserved up front.

  6. Speculative decoding: Here, a small model (speculator) proposes multiple tokens. After that, the main model verifies them in one forward pass.

Over to you: what is missing from this list?


[Webinar] How to stop babysitting your agents (Sponsored)

Agents can generate code. Getting it right for your system, team conventions, and past decisions is the hard part. You end up wasting time and tokens in the correction loops.

More MCPs, rules, and bigger context windows give agents access to information, but not understanding. The teams pulling ahead have a context layer to give agents exactly what they need for the task at hand.

Join us for a FREE webinar on Oct 7 to see:

  • Where teams get stuck on the AI maturity curve and why common fixes fall short

  • How a context layer solves for quality, efficiency, and cost

  • Live demo: the same coding task with and without a context layer

If you want to maximize the value you get from AI agents, this one is worth your time.

Register Now


The Anatomy of an AI Agent

An AI agent can be thought of as a simple While-loop.

It uses an LLM to select an action, executes that action, evaluates the result, and repeats the process until the task is complete. Let’s take a closer look at each of these components:

  • Brain: The LLM is the core. It reads the situation, thinks, and decides what to do next. The big shift from chatbot to agent: the model isn’t writing text anymore, it’s making choices.

  • Planning: Hard tasks need more than one step. Agents break them down using methods like Chain of Thought (think step by step), Tree of Thoughts (try options, pick the best), or
    Reflexion (learn from mistakes and retry). Planning turns a fuzzy goal into clear actions.

  • Tools: An LLM without tools is a brain in a jar. Tools are functions the model can call, like web search, code execution, APIs, files, or browsers (often using the MCP standard). The model requests a tool, the system runs it, and the result comes back.

  • Memory: Without memory, every turn starts from zero. Short-term memory is the context window. Long-term memory lives in vector stores, files, and knowledge bases. When the window fills up, agents summarize old turns and carry the summary forward.

  • Loop: All four pieces work together in a cycle. The agent looks at the current state, decides what to do, uses a tool, sees the result, and repeats. It keeps going until it gives a final answer.

  • Guardrails: Not strictly anatomy, but important. Sandboxing, human checks, token limits, output validation, and scope limits keep autonomy from turning into expensive chaos. The more autonomy you give, the more these matter.

Over to you: when you build an agent, which of these five takes the most work to get right?


How do you know if your AI app actually works?

You evaluate it. But most teams skip this step (or do it wrong) because “eval” feels vague. It’s not.

Every good eval is a 3-step recipe.

Step 1: Pick a task. AI systems have different capabilities and dimensions to evaluate. For LLMs, it can be safety or math capability, in RAGs it can be grounding and retrieval, Pick one.

Step 2: Collect eval data. For every task, gather inputs paired with the right answer or expected behavior. A safety set pairs risky prompts with “refuse.”

Step 3: Develop a grader. How do you decide if the output is good?

  • Use code-based graders (if/else, unit tests) for things with a clear correct answer and patch passing unit-tests.

  • Use model-based graders (LLM-as-judge) for subjective tasks like safety.

  • Use human graders for edge cases and anything where nuance matters more than throughput.

Most production evals combine all three. Code-based for what’s cheap to check. Model-based for scale. Human-based for what matters most.

Over to you: what’s the hardest thing about your task to grade, and which grader type do you use for it?


🚀 New Course: AI Evals in Practice Starts on Oct. 7

ByteByteGo has teamed up with Manjeet Singh, Senior Director at Salesforce, to bring you a live and hands-on course on building reliable evaluation systems for production AI agents.

Check it out Now

You’ll learn how to:

  • Design evals for quality, safety, reliability, cost, and latency

  • Build and validate LLM-as-a-Judge systems

  • Red-team agents for prompt injection and jailbreaks

  • Create meaningful eval datasets from real and synthetic data

  • Evaluate tool use, RAG, multi-step execution, and multi-agent handoffs

  • Run evals in CI/CD and production to catch regressions and drift

  • Turn failures into permanent regression tests

If you’re building AI agents and asking, “How do I know this is actually ready to ship?” — this course is for you.

Check it out Now

Why State is the Hardest Thing in Software Design

2026-10-01 23:31:21

State is the information a system retains that affects its functionality in some way. For example, a website might keep track of users who have logged in. A document editor remembers the current text. A background worker logs which tasks it has completed and which ones are pending. A database stores the records written to it for as long as needed.

Any useful software usually needs to manage state to carry out its intended functionality. The difficulty is keeping the state correct even as requests overlap, servers are added, machines fail, and software changes are deployed.

Often, the common advice is to “make the application stateless”. But this advice needs explanation. It usually means making particular application servers easy to replace while putting their important state elsewhere. It doesn’t signify a complete absence of state. Understanding such a setup requires looking at why state exists, who owns it, and what happens when it becomes unavailable.

In this article, we will look at why state is the hardest thing in software development and the multiple strategies that developers can use to manage it.

What Is State?

Read more

How DoorDash Built a Toolbox for AI Agents

2026-09-30 23:31:13

CodeAF: a new open-source factory on the Pareto frontier (Sponsored)

CodeAF is a new open-source software factory that sits on the Pareto frontier of cost, speed and quality. On DeepSWE it solved nearly 4× as many real GitHub issues as Claude Code on the same open model, and matched the official leaderboard result at half the cost.

It is built from the ground up for the next era of coding, where you stop chatting with agents and start directing them. Built for open models like DeepSeek, Qwen, GLM and Kimi, it gives you frontier-grade coding without locking you into a closed model, and it still works with any provider you choose.

Try CodeAF on GitHub


The utility of AI agents increases when they can take real actions in real systems. This involves the use of tools. While MCP made it easier to expose these tools, there were a lot of other concerns that had to be handled in order to make it work at an enterprise level.

DoorDash built a shared Agent Gateway to control how AI agents discover and use tools. The gateway brings together several responsibilities: checking permissions, managing credentials, choosing which tools an agent can see, forwarding requests, and recording what happened.

In this article, we will look at how the DoorDash engineering team built this gateway and the decisions they made. Here’s what we will cover:

  • Why does an AI agent need tools

  • Why MCP is not enough

  • The core components of the gateway

  • Verifying who is calling and what they can do

  • Why identifying the caller is different from supplying credentials

  • How a user connects an account during a tool call

  • Why agents should see a tool catalog

  • What happens during discovery and execution

  • How DoorDash made the platform easy to adopt

Disclaimer: This post is based on publicly shared details from various sources. References at the end. Please comment if you notice any inaccuracies.

Why an AI Agent Needs Tools

A language model can write a super-detailed explanation of how to investigate a software problem, but it doesn’t automatically have access to a company’s code repositories, incident reports, or production logs. We need to provide access to those systems through software. From the perspective of the model, these systems are like tools.

An AI agent is an application that uses a language model to decide which steps and tools to use while working on a particular task. A tool is a specific capability that the application makes available to the model. For example, searching documentation, reading a support ticket, and opening a pull request.

When the model selects a tool, the surrounding application uses the tool to execute the required operation and supplies the result back to the language model. The model can then use that result to decide what to do next. For example, an agent investigating a failed software build might retrieve the build logs, inspect relevant code, and use what it finds to explain the failure.

Some tools only read information. Others can also change something in a real system, such as updating a ticket or creating a pull request. This makes tool access a very important part of an agent’s design. It determines what the agent can learn and what it can eventually do with that learning. At DoorDash, these capabilities come from many places, including internal services, engineering systems, documentation platforms, and third-party software.

Why MCP Is Not Enough

Different systems or tools normally expose different interfaces. An API is a way for one program to request information or actions from another program. Without a shared approach, an AI agent application has to accommodate the particular interfaces of every tool it uses.

The Model Context Protocol (MCP) provides a common way to describe, discover, and invoke capabilities. An MCP server exposes tools, and an MCP client communicates with that server. The agent application uses an MCP client to access the tools.

This is largely facilitated by two main operations:

  • The first is tools/list, which discovers the tools available from a server. The returned catalog gives the agent information about the available capabilities, including their names, descriptions, and expected inputs.

  • The second is tools/call, which requests execution of a particular tool with the supplied inputs.

For example, a server might expose a tool for looking up an issue in the issue tracker. A discovery service tells the agent that this capability exists and how to request it. Invocation asks the server to perform a specific lookup.

Such a shared interface makes integration simple. However, DoorDash still had to decide which agents can use which tools, whose account an action should use, and how access should be monitored. A coding agent might need access to GitHub, Jira, code search, documentation, and systems that show production behavior. Each connection introduces decisions about permissions, credentials, available operations, and monitoring.

This makes the problem of tool access not just a matter of connection. Also, there are a bunch of company-specific rules that come into the picture here. MCP’s common interface doesn’t establish those company-specific rules. If every team implements these responsibilities independently, the same work gets repeated across many agents and servers. For example, one team might build an OAuth connection flow, another might develop its own secret handling, and another creates a different way to record tool usage. It becomes quite difficult to maintain consistent behavior.

DoorDash organizes this problem into three separate concerns:

  • Access: Deals with the identity and permissions behind a request. The platform needs to know who is calling, what that caller is allowed to do, and which credentials the downstream system requires. For example, an agent acting on behalf of an individual employee may need a different credential arrangement from an automated process running for an entire team.

  • Tool-surface Curation: This concern handles the capabilities presented to the agent. A downstream server might offer hundreds of tools, while a particular workflow only needs a small subset. DoorDash needed a way to select the relevant, approved tools instead of exposing the entire catalog.

  • Operations: This concerns what happens when these integrations run in production. Teams need to understand traffic, failures, response times, usage, and costs. They also need controls that prevent excessive requests from overwhelming downstream systems.

The Agent Gateway brings all of these concerns and responsibilities into a shared platform. Let’s look at that in more detail.

The Core Components of the Agent Gateway

A gateway is a service through which requests pass before reaching their destination.

In DoorDash’s design, an agent sends its MCP requests to the Agent Gateway, which controls access and forwards approved requests to the appropriate downstream MCP server. Here, downstream simply means the system that receives a request after the gateway forwards it.

There are two core components involved here: a proxy and a registry.

The proxy handles requests as they arrive. It verifies the caller, checks permissions, applies rate limits, attaches the required credentials, and forwards the request. It also produces information used for monitoring and auditing. Since it handles the actual request traffic, it is considered part of the data plane.

The registry stores the information needed to govern the incoming traffic. This includes details about the registered agents and MCP servers, ownership details, connection settings, authentication modes, policies, discovered tool catalogs, and configurations that determine which tools are exposed.

The registry is the source of truth for the control plane, which manages how the system should operate. The engineering teams use a management interface and API to configure the gateway, while the proxy applies the resulting configuration to requests.

This division between the two components separates the work of configuring access from the work of handling traffic. A tool owner can register a capability and attach policy through the control plane. The proxy then enforces that policy when the AI agents need to use the capability.

DoorDash also separates internal workflows and external-facing use cases into different proxy planes. They share libraries and registry concepts. But they maintain separate trust boundaries. Think of it as a common governance approach without implying that every request passes through one physical proxy instance

Verifying Who is Calling and What They Can Do

Two closely related concepts play a key role in the Agent Gateway built by the DoorDash engineering team. These concepts are authentication and authorization.

Authentication establishes the identity of the caller making the request. A request might come from a user, a software service, or an agent acting on behalf of a user. Authorization determines what that identified caller is allowed to do.

These checks need more information than simply deciding whether someone can access an entire MCP server. A server may expose both harmless lookup operations and powerful administrative actions. The distinction between the two should be clear. For example, permission to search a repository should not automatically imply permission to use every operation the server provides.

DoorDash’s agent gateway can evaluate permissions involving several pieces of context. It can consider the agent, the user, the requested tool, and the environment in which the request occurs. It can also distinguish access to read-only capabilities from access to capabilities that modify data. For example, a policy might allow an agent to read issue details on behalf of a user while preventing that agent from performing administrative operations. Another workflow might be allowed to run under a team identity, while a user-specific action requires the individual user’s authorization.

The agent gateway carries the identity of the caller through the rest of the request process, including routing and recording usage. That makes it possible to attribute a downstream action to the context that caused it.

Centralizing these decisions also gives DoorDash a shared place to change or revoke access if needed. Agent developers don’t have to encode security decisions in prompts. Also, each tool owner doesn’t have to rebuild the same gateway-level authorization logic.

Why Identifying the Caller is Different from Supplying Credentials

Even after the agent gateway knows who is making a request and whether it is allowed, the downstream system may require its own credentials.

A credential is something used to establish identity or access, such as an API key or an access token. The identity used to enter DoorDash’s gateway and the credential required by a third-party service are separate parts of the request.

For instance, the gateway may recognize an employee through the internal identity system. However, a documentation provider may still require that employee’s authorization token before it allows access to their documents. There are four types of credential arrangements that are possible:

  • Internal Service Identity: This is an identity verified by the DoorDash infrastructure. It forwards verified caller context to internal services.

  • Gateway-held Token: This is a vendor or service token kept in gateway secret storage. It adds the token to the downstream request without giving it to the agent.

  • Per-user OAuth: This is for access that a particular user has granted to an external service. The gateway stores the grant securely, supplies the user’s token, and refreshes it when needed.

  • Service Principal: This is a non-personal identity used by software or team automation. The gateway obtains or issues short-lived credentials through a gateway-managed team identity.

Credential injection means adding the appropriate credential to the outgoing request before forwarding it. The agent asks to use a tool. The gateway handles the credential needed to access the system behind that tool.

The design created by the DoorDash engineering team keeps raw downstream credentials, including vendor keys and OAuth refresh tokens, out of the agents. This gives the platform a central place to manage, rotate, audit, and revoke those credentials.

How a User Connects an Account During a Tool Call

Some tools need to act within a particular user’s account. For example, reading that user’s documents or updating a ticket as that user requires appropriate user authorization.

OAuth is a mechanism through which a user authorizes an application to access a service with specified permissions. The application receives tokens that it can use for the authorized access.

The access token is used when making requests. A refresh token, when issued, allows the application to obtain a replacement access token when the current one expires.

Before the agent gateway, the product teams within DoorDash tended to build their own connection flows and token storage. However, the gateway centralizes this work.

When a call needs authorization that is not yet available, the gateway starts the provider’s OAuth flow. After the user authorizes access, the gateway stores the resulting tokens in encrypted form and supplies the appropriate token on future calls.

For clients that support MCP elicitation, the gateway can ask the client to present a connection prompt while keeping the original tool request open. The user completes authorization in a browser. The gateway then obtains and stores the token, resumes the original request with that token, and returns the result from the tool.

To clarify things here, elicitation is the protocol mechanism used to request the user interaction needed to continue the task. The user completes the provider’s authorization flow. However, the agent doesn’t receive the raw credential.

For clients without elicitation support, the gateway returns a structured response explaining that authorization is required, along with a connection URL. This gives the client a clear way to recover, even though it cannot use the same pause-and-resume interaction.

Why Agents Should See a Selected Tool Catalog

Permissions determine which actions are allowed, but the tool surface shown to an agent also affects its behavior.

Tool surface means the set of tools presented to an agent, including their names, descriptions, and organization. A server’s complete catalog may contain administrative functions, billing operations, destructive actions, and features unrelated to the agent’s task. Presenting all of those capabilities creates unnecessary choices. The model has more descriptions to interpret and more potentially similar tools to distinguish. The approach used by the DoorDash engineering team is to expose smaller, relevant catalogs that match the work being performed.

There are two mechanisms behind this: bundles and filters.

A bundle combines tools from multiple MCP servers behind one logical MCP endpoint. An endpoint is the address to which the client sends requests. For example, a developer-tools bundle can provide selected GitHub, Jira, observability, code-search, and documentation capabilities through one gateway address. The tools still run through their respective downstream servers. However, the bundle gives the agent one organized interface through which to discover and use them.

Filters determine which tools appear in that interface. Their decisions can depend on the bundle, agent, user group, environment, or intended audience. A provider might expose both issue lookup and administrative operations, while a particular bundle includes only the approved issue-related capabilities.

The agent gateway can also give tools stable names and clearer descriptions. When names from different servers overlap, namespacing or aliases can help distinguish them. Namespacing means adding identifying context to a name, such as identifying the service to which a tool belongs.

What Happens During Discovery and Execution

Discovery and execution are related, but they happen at different points and require their own checks and validations.

During discovery, the agent sends tools/list to a bundle endpoint. The gateway gathers tool information from the servers included in that bundle. This is known as fan-out, meaning that one incoming request leads to requests across multiple downstream servers.

The gateway then applies authorization and filtering rules. It combines the approved tools into a coherent catalog, adjusts names where necessary, and returns the result to the agent. The agent might therefore receive a catalog containing repository tools from one server, issue tools from another, and log-query tools from a third. It interacts with the gateway’s combined interface instead of managing each server’s setup separately.

When the agent chooses a tool, it sends a tools/call request to the gateway. The gateway checks policy again, applies the relevant rate limits, determines the correct downstream server, and supplies the required credential. It then forwards the request and returns the result.

A tool appearing in a discovery response doesn’t remove the need to authorize its execution.

In other words, discovery controls what is presented to the agent. Invocation controls whether a particular request is allowed to proceed.

This gives the gateway responsibility for both the available catalog and the actual use of its tools. It also means the agent doesn’t need to implement downstream routing. Neither does it have to select credentials for each provider.

Understanding What Happened After a Request

Once many agents have used many tools, the teams need to understand more about the resulting activity. This comes under the area of observability, which means the ability to understand a running system through the information it produces.

The gateway acts as a useful component for this because all tool requests pass through it. It can record a structured event for each call, with consistent fields describing the server, tool, bundle, owner, and identities involved. For reference, a structured event is a record with defined fields that can be searched or aggregated. Along with identity and ownership, an event can also contain the authorization result, request status, error source, timing information, and request and response sizes.

This helps answer questions such as which agent generated a request, which tool failed, and which team owns the affected capability.

The gateway also produces metrics, which summarize behavior across requests. This includes request counts, tool response times, authorization decisions, OAuth refresh outcomes, rate-limit decisions, and downstream failures.

These different forms of information serve different groups within the enterprise. Security teams can audit access, platform teams can investigate unusually active agents, and tool owners can understand adoption and failures.

Making the Platform Easy for Teams to Adopt

A centralized gateway is useful only if teams can use it without excessive effort. Therefore, DoorDash makes registration and configuration for the gateway available through a self-service management interface and API.

Here’s how the process works:

  • A team starts by registering its MCP server. The platform discovers the server’s raw catalog through tools/list, after which the tools intended for exposure are selected and approved.

  • The team attaches the authentication mode, ownership information, and access policy.

  • Approved tools are then added to one or more bundles. Once agents begin using them, teams can inspect traffic, response times, failures, authorization decisions, and cost information.

As we can see, registration is the beginning of onboarding. Discovering that a server offers a capability doesn’t automatically mean that every agent receives access to it. Ownership information is also part of making the system operationally useful. When a capability fails or needs a policy change, the platform needs a clear association between that capability and the team responsible for it.

The gateway has now become the default path for agent-tool access across DoorDash engineering teams. DoorDash has reported more than 200 registered MCP servers, more than 30 agents and services used by thousands of employees, and millions of tool calls each week.

What DoorDash Plans to Improve Next

DoorDash is making further investments in the gateway.

One planned improvement is stronger agent identity and user delegation. Delegation means allowing an agent to perform an action on someone else’s behalf within defined permissions.

DoorDash’s target model gives each agent a cryptographic identity and uses short-lived delegated credentials scoped to the user, agent, task, and target tool. This would make it possible to determine the connection between the user requesting work, the agent performing it, and the specific action being authorized.

Another planned improvement is dynamic tool discovery. Existing bundles already narrow the tools available to an agent. Dynamic discovery would narrow that selection further using the current task, access policy, and usage signals, so the agent receives the capabilities most likely to help with its present work.

There are also plans to improve evaluation of tool quality and security, identifying risky descriptions, detecting secrets or personally identifiable information in errors, simplifying server creation and registration, and producing redacted tool-call event streams.

These investments extend the same architectural approach. As agents gain access to more systems, the platform needs precise ways to establish who is acting, expose useful capabilities, authorize specific actions, and make the resulting activity understandable.

References