Skip to main content

Command Palette

Search for a command to run...

🧠 Understanding Copilot Agents in VS Code

Updated
•7 min read•View as Markdown
P
Senior Software Engineer who likes to code, write and climb trees.

Note: VS Code Copilot’s new Agent Mode isn’t limited to Claude. It can work with any supported model that implements the same protocol. I’m using Claude 3.7 as an example here because it’s currently one of the best at reasoning and tool use inside copilot.

šŸ¤” What Changed in VS Code Copilot?

Earlier, we could only chat with Claude or GPT inside copilot.
You’d ask a question, it would reply, and that was it.

Now there’s a new thing called Agent Mode.

It’s not just about chatting anymore. It’s about having an assistant that can actually act:
read files, edit code, search your repo, and plan multi-step tasks while keeping context.


🧩 Ask Mode vs Agent Mode (The Way I Understood It)

🧩 Ask Mode vs Agent Mode (Simplified Explanation)

In Ask Mode, you talk to Claude like a normal chatbot.
It seems like it remembers what you said earlier, but technically it doesn’t have real memory.
Each time you send a new message, the app (like VS Code copilot or the Claude web app) resends your past conversation so Claude can respond in context.
That’s why it feels continuous, even though every message is technically a fresh API call.

In Agent Mode, things change.
When you start an agent session, Claude’s cloud runtime creates a live, temporary session that stays active for a while (usually a few hours).
This live session holds context, tool call results, and task progress.
Your VS Code just keeps the connection open — it doesn’t store or manage the state itself.
That’s why Claude in Agent Mode can remember what it just did, plan next steps, and interact with your local tools until the session ends.

Ask Mode Agent Mode
How it works Each message is sent as a new request, but the app resends the full chat history so it feels continuous. Starts a live, temporary session in Claude’s cloud runtime that stays active for the duration of your task.
Context handling Context is rebuilt every time from past messages. Context is maintained live in the cloud session.
What it can do Text-only reasoning, no real actions. Can use local tools (search, edit, run) through the MCP client.
Memory Recreated for every message. Persistent while the live session runs.
Where the state lives Nowhere, each call is independent. In Claude’s cloud runtime (not stored locally).
Best for Quick Q&A or lightweight reasoning. Multi-step automation, code editing, and planning.

āš™ļø What’s Happening Technically Behind the Scenes?

When you start an agent like Claude in VS Code, three main parts work together:

  • Claude (cloud) → The model that plans what to do next

  • MCP Client (VS Code) → The bridge between the cloud and your computer

  • MCP Servers (local) → Small executors that perform the actual work, like searching, editing, or running commands

Your setup looks like this:

Claude (cloud)
⇅ JSON over HTTPS
MCP Client (VS Code)
⇅ Local IPC / JSON-RPC
MCP Servers (your machine)

🧩 What Is an MCP Server?

An MCP Server is basically a local mini API that provides tools for the model to use.

Examples of built-in servers:

  • Filesystem MCP → read, edit, or search files

  • Git MCP → get diffs or commit info

  • Terminal MCP → run shell commands

Each declares its available tools through a simple manifest, like this:

{
  "name": "filesystem",
  "tools": ["read_file", "edit_file", "search_code"]
}

So when Claude wants to search your code, it doesn’t do that directly.
It sends a structured JSON request like this:

{
  "type": "call_tool",
  "name": "search_code",
  "arguments": { "pattern": "ButtonPrimary" }
}

The MCP Client (Vscode copilot in this case) routes this request to the right MCP Server (in this case, filesystem),
and the server executes it locally, similar to running a grep search.
Then the results are sent back to Claude for analysis.


šŸ’­ Does Claude Remember All This?

Yes, but only to a certain point.

Layer Keeps Context? Notes
Claude (cloud runtime) āœ… Yes Remembers tool calls, results, and messages while the session is active
MCP Client (VS Code) āŒ No Only routes messages
MCP Servers (local) āŒ No Stateless executors

So Claude is the only part that holds memory, and even that is temporary. It lasts only while the agent session is active.

When you stop the session, that memory is gone.


šŸ• Then How Come I See My Old Agent Chats?

Good question.

That happens because VS Code copilot stores the visible chat locally.
When you reopen an agent, VS Code copilot re-sends that chat history to Claude’s API.

So it looks like Claude remembers, but it doesn’t. It’s just reading what VS Code copilot replayed.


🧮 About the 200k Token Context

Claude 3.5 Sonnet supports up to 200k tokens of live context.
That includes everything:

  • Your messages and Claude’s replies

  • Tool call inputs and outputs

  • Code snippets and search results

When you reach the limit, Claude starts summarizing older parts to stay within it.
So the longer your chat or the larger your snippets, the less room you have left for new reasoning.


āœ… How To Avoid Losing Context (Claude’s Truncation)

Here are a few simple habits that help:

  1. Keep chats focused.
    Don’t send entire files; share smaller diffs or snippets.

  2. Summarize regularly.
    Ask Claude to condense what’s been done:
    ā€œSummarize what we’ve done so far so we can continue with a smaller context.ā€

  3. Split tasks into smaller sessions.
    One session for refactoring, another for testing, and so on.

  4. Save checkpoints.
    When you reach a good point, save that plan or code change into a markdown file instead of keeping it only in chat.

  5. Limit verbose outputs.
    You can say things like: ā€œShow only filenames, not all lines.ā€


⚔ Example: Refactoring a Component with Claude Agent

Let’s say you tell Claude:

ā€œRename all uses of ButtonPrimary to ButtonMain.ā€

Here’s what actually happens behind the scenes:

1ļøāƒ£ Claude (cloud)
Plans the next step: ā€œFind all ButtonPrimary imports.ā€
Sends a search_code tool call.

2ļøāƒ£ MCP Client (VS Code copilot)
Routes that call to the filesystem MCP server.

3ļøāƒ£ MCP Server (local)
Runs a local search, like grep or ripgrep.
Returns the results as JSON.

4ļøāƒ£ MCP Client
Forwards those results back to Claude.

5ļøāƒ£ Claude (cloud)
Analyzes matches, plans edits, and sends another tool call (edit_file).

6ļøāƒ£ MCP Server
Applies the edits and confirms success or failure.

So:

  • Thinking and planning happen in the cloud

  • Execution happens locally on your machine


🧠 In Short

  • Claude is the brain (cloud)

  • MCP Client is the messenger (VS Code copilot )

  • MCP Servers are the hands (your local tools)


šŸ” Summary

Concept Role Lives Where Persists?
Claude model Brain (planner) Cloud Until session ends
MCP client Router VS Code No
MCP server Executor Local No
Context window Short-term memory Cloud 200k tokens
Chat history Local copy VS Code Yes (text only)

That’s it.
That’s how Claude Agents in VS Code move from being a simple chat feature to a context-aware, locally capable assistant that can think, plan, and act almost like a real developer sitting beside you.