š§ Understanding Copilot Agents in VS Code
Note: VS Code Copilotās new Agent Mode isnāt limited to Claude. It can work with any supported model that implements the same protocol. Iām using Claude 3.7 as an example here because itās currently one of the best at reasoning and tool use inside copilot.
š¤ What Changed in VS Code Copilot?
Earlier, we could only chat with Claude or GPT inside copilot.
Youād ask a question, it would reply, and that was it.
Now thereās a new thing called Agent Mode.
Itās not just about chatting anymore. Itās about having an assistant that can actually act:
read files, edit code, search your repo, and plan multi-step tasks while keeping context.
š§© Ask Mode vs Agent Mode (The Way I Understood It)
š§© Ask Mode vs Agent Mode (Simplified Explanation)
In Ask Mode, you talk to Claude like a normal chatbot.
It seems like it remembers what you said earlier, but technically it doesnāt have real memory.
Each time you send a new message, the app (like VS Code copilot or the Claude web app) resends your past conversation so Claude can respond in context.
Thatās why it feels continuous, even though every message is technically a fresh API call.
In Agent Mode, things change.
When you start an agent session, Claudeās cloud runtime creates a live, temporary session that stays active for a while (usually a few hours).
This live session holds context, tool call results, and task progress.
Your VS Code just keeps the connection open ā it doesnāt store or manage the state itself.
Thatās why Claude in Agent Mode can remember what it just did, plan next steps, and interact with your local tools until the session ends.
| Ask Mode | Agent Mode | |
|---|---|---|
| How it works | Each message is sent as a new request, but the app resends the full chat history so it feels continuous. | Starts a live, temporary session in Claudeās cloud runtime that stays active for the duration of your task. |
| Context handling | Context is rebuilt every time from past messages. | Context is maintained live in the cloud session. |
| What it can do | Text-only reasoning, no real actions. | Can use local tools (search, edit, run) through the MCP client. |
| Memory | Recreated for every message. | Persistent while the live session runs. |
| Where the state lives | Nowhere, each call is independent. | In Claudeās cloud runtime (not stored locally). |
| Best for | Quick Q&A or lightweight reasoning. | Multi-step automation, code editing, and planning. |
āļø Whatās Happening Technically Behind the Scenes?
When you start an agent like Claude in VS Code, three main parts work together:
Claude (cloud) ā The model that plans what to do next
MCP Client (VS Code) ā The bridge between the cloud and your computer
MCP Servers (local) ā Small executors that perform the actual work, like searching, editing, or running commands
Your setup looks like this:
Claude (cloud)
ā
JSON over HTTPS
MCP Client (VS Code)
ā
Local IPC / JSON-RPC
MCP Servers (your machine)
š§© What Is an MCP Server?
An MCP Server is basically a local mini API that provides tools for the model to use.
Examples of built-in servers:
Filesystem MCP ā read, edit, or search files
Git MCP ā get diffs or commit info
Terminal MCP ā run shell commands
Each declares its available tools through a simple manifest, like this:
{
"name": "filesystem",
"tools": ["read_file", "edit_file", "search_code"]
}
So when Claude wants to search your code, it doesnāt do that directly.
It sends a structured JSON request like this:
{
"type": "call_tool",
"name": "search_code",
"arguments": { "pattern": "ButtonPrimary" }
}
The MCP Client (Vscode copilot in this case) routes this request to the right MCP Server (in this case, filesystem),
and the server executes it locally, similar to running a grep search.
Then the results are sent back to Claude for analysis.
š Does Claude Remember All This?
Yes, but only to a certain point.
| Layer | Keeps Context? | Notes |
|---|---|---|
| Claude (cloud runtime) | ā Yes | Remembers tool calls, results, and messages while the session is active |
| MCP Client (VS Code) | ā No | Only routes messages |
| MCP Servers (local) | ā No | Stateless executors |
So Claude is the only part that holds memory, and even that is temporary. It lasts only while the agent session is active.
When you stop the session, that memory is gone.
š Then How Come I See My Old Agent Chats?
Good question.
That happens because VS Code copilot stores the visible chat locally.
When you reopen an agent, VS Code copilot re-sends that chat history to Claudeās API.
So it looks like Claude remembers, but it doesnāt. Itās just reading what VS Code copilot replayed.
š§® About the 200k Token Context
Claude 3.5 Sonnet supports up to 200k tokens of live context.
That includes everything:
Your messages and Claudeās replies
Tool call inputs and outputs
Code snippets and search results
When you reach the limit, Claude starts summarizing older parts to stay within it.
So the longer your chat or the larger your snippets, the less room you have left for new reasoning.
ā How To Avoid Losing Context (Claudeās Truncation)
Here are a few simple habits that help:
Keep chats focused.
Donāt send entire files; share smaller diffs or snippets.Summarize regularly.
Ask Claude to condense whatās been done:
āSummarize what weāve done so far so we can continue with a smaller context.āSplit tasks into smaller sessions.
One session for refactoring, another for testing, and so on.Save checkpoints.
When you reach a good point, save that plan or code change into a markdown file instead of keeping it only in chat.Limit verbose outputs.
You can say things like: āShow only filenames, not all lines.ā
ā” Example: Refactoring a Component with Claude Agent
Letās say you tell Claude:
āRename all uses of ButtonPrimary to ButtonMain.ā
Hereās what actually happens behind the scenes:
1ļøā£ Claude (cloud)
Plans the next step: āFind all ButtonPrimary imports.ā
Sends a search_code tool call.
2ļøā£ MCP Client (VS Code copilot)
Routes that call to the filesystem MCP server.
3ļøā£ MCP Server (local)
Runs a local search, like grep or ripgrep.
Returns the results as JSON.
4ļøā£ MCP Client
Forwards those results back to Claude.
5ļøā£ Claude (cloud)
Analyzes matches, plans edits, and sends another tool call (edit_file).
6ļøā£ MCP Server
Applies the edits and confirms success or failure.
So:
Thinking and planning happen in the cloud
Execution happens locally on your machine
š§ In Short
Claude is the brain (cloud)
MCP Client is the messenger (VS Code copilot )
MCP Servers are the hands (your local tools)
š Summary
| Concept | Role | Lives Where | Persists? |
|---|---|---|---|
| Claude model | Brain (planner) | Cloud | Until session ends |
| MCP client | Router | VS Code | No |
| MCP server | Executor | Local | No |
| Context window | Short-term memory | Cloud | 200k tokens |
| Chat history | Local copy | VS Code | Yes (text only) |
Thatās it.
Thatās how Claude Agents in VS Code move from being a simple chat feature to a context-aware, locally capable assistant that can think, plan, and act almost like a real developer sitting beside you.


