How to Use DeepSeek in VS Code: The $2/Month AI Coding Assistant Setup Guide

Coding Liquids blog cover featuring Sagnik Bhattacharya for How to Use DeepSeek in VS Code, showing VS Code editor with DeepSeek AI coding assistant integration.
Coding Liquids blog cover featuring Sagnik Bhattacharya for How to Use DeepSeek in VS Code, showing VS Code editor with DeepSeek AI coding assistant integration.

GitHub Copilot Pro costs $10 per month, and a Business seat is $19. It is the industry standard for AI-assisted coding, and for good reason — it is fast, accurate, and deeply integrated into VS Code. But what if I told you there is an alternative that delivers frontier-class coding quality for roughly $2 per month? That alternative is DeepSeek, and it has quietly become one of the most compelling budget AI coding assistants available in 2026.

I teach Flutter and Excel with AI — explore my courses if you want structured learning.

DeepSeek, developed by the Chinese AI lab of the same name, has made waves with its V3, R1, and V4 model families. As of 24 September 2026 the API offers two models. The everyday one is DeepSeek V4.1 Flash, released on 10 September 2026 under the API name deepseek-flash: a 552B-parameter mixture-of-experts model with a new causal encoder–decoder design that activates only 8B parameters while reading your prompt and 16B while writing the answer, with native image input and open weights on Hugging Face. Above it sits DeepSeek V4 Pro (deepseek-v4-pro), the larger reasoning tier that went GA on 13 August 2026. Both expose a 1M-token context window and up to 384K output tokens. The pricing is what makes DeepSeek remarkable: Flash costs $0.15–0.30 per million input tokens and $0.60–1.20 per million output tokens (off-peak is half the peak rate), a fraction of what OpenAI or Anthropic charge for comparable models. For the average developer writing code in VS Code, that translates to roughly $1.50 to $3 per month in real-world usage. If you prefer complete privacy and zero cost, you can also run DeepSeek Coder V2 locally through Ollama.

Note on legacy model names: The old deepseek-chat and deepseek-reasoner aliases were retired at 15:59 UTC on 24 July 2026 — any config still using them now gets an error back instead of a completion, so fix those first. deepseek-v4-flash (and the experimental deepseek-v4-flash-vision-exp) still resolve for now, but V4 Flash itself has been retired and those names are temporarily routed to V4.1 Flash at Flash pricing. Swap them for deepseek-flash so you are not caught out when that routing ends. The deepseek-coder API model name was folded into deepseek-chat back when V2.5 launched in 2024. Use deepseek-flash and deepseek-v4-pro in any new configuration.

Follow me on Instagram@sagnikteaches

I have spent the past few weeks setting up DeepSeek in VS Code using multiple methods, testing it across real coding tasks, and tracking my actual API costs. This guide walks through every step — from getting your API key to configuring autocomplete to understanding exactly what you are trading off compared to Copilot and Claude.

Connect on LinkedInSagnik Bhattacharya

Prerequisites

Before you begin, make sure you have the following ready:

Subscribe on YouTube@codingliquids
  1. VS Code installed. Any recent stable release works. Keep it updated for the best extension compatibility.
  2. A DeepSeek API key (for the cloud method). You will create an account at platform.deepseek.com and generate an API key. DeepSeek requires a small initial top-up (as little as $2) to activate API access.
  3. Ollama installed (for the local method). Download it from ollama.com if you want to run DeepSeek Coder entirely on your machine with no API costs.
  4. The Continue extension. This is the open-source VS Code extension we will use to connect DeepSeek to your editor. It supports both API-based and local models, and provides chat, inline editing, and tab-autocomplete.

You do not need a powerful GPU for the API method — the model runs on DeepSeek's servers. For the local Ollama method, you will need at least 8GB of VRAM for the smaller DeepSeek Coder models.

Method 1: DeepSeek API + Continue Extension

This is the recommended method for most developers. You get access to DeepSeek's most powerful models (Flash and V4 Pro) without any hardware requirements, and the cost is negligible compared to Copilot.

Step 1: Get Your DeepSeek API Key

  1. Go to platform.deepseek.com and create an account.
  2. Navigate to the API Keys section in your dashboard.
  3. Click "Create new API key" and copy the key. Store it somewhere safe — you will not be able to see it again.
  4. Top up your account with the minimum amount (typically $2). This balance will last most developers several weeks to a month depending on usage.

DeepSeek's API is OpenAI-compatible, which means any tool that works with the OpenAI API format can connect to DeepSeek by simply changing the base URL and API key. This is why the Continue extension works seamlessly with it.

Step 2: Install and Configure Continue

  1. Open VS Code and go to the Extensions panel (Ctrl+Shift+X).
  2. Search for "Continue" and install the extension by Continue.dev.
  3. After installation, click the Continue icon in the sidebar.
  4. Open the Continue configuration file. The current Continue extension uses YAML by default at ~/.continue/config.yaml. You can open it via the gear icon in the Continue panel, or with the command palette (Ctrl+Shift+P → "Continue: Open config.yaml"). The legacy config.json format still loads if present, but new installs are YAML.

Add the following to your config.yaml:

name: My Continue Config
version: 1.0.0
schema: v1
models:
  - name: DeepSeek Flash (Fast)
    provider: deepseek
    model: deepseek-flash
    apiKey: YOUR_DEEPSEEK_API_KEY
    roles:
      - chat
      - edit
      - apply
    requestOptions:
      extraBodyProperties:
        thinking:
          type: disabled
  - name: DeepSeek Flash (Thinking)
    provider: deepseek
    model: deepseek-flash
    apiKey: YOUR_DEEPSEEK_API_KEY
    roles:
      - chat
  - name: DeepSeek V4 Pro
    provider: deepseek
    model: deepseek-v4-pro
    apiKey: YOUR_DEEPSEEK_API_KEY
    roles:
      - chat
      - edit
  - name: DeepSeek Autocomplete
    provider: deepseek
    model: deepseek-flash
    apiKey: YOUR_DEEPSEEK_API_KEY
    roles:
      - autocomplete
    autocompleteOptions:
      debounceDelay: 300
      maxPromptTokens: 1024

Replace YOUR_DEEPSEEK_API_KEY with the key you copied earlier. The configuration above gives you two views of the same Flash model — a "Fast" entry with thinking switched off for instant edits and quick questions, and a "Thinking" entry that leaves DeepSeek's default reasoning on for harder problems — plus V4 Pro for the genuinely difficult work, and Flash again as the autocomplete model, because latency matters more than depth for tab suggestions. The extraBodyProperties block is how you reach DeepSeek's thinking request parameter from Continue: the API turns thinking on by default (at high effort), and Continue's DeepSeek provider has no dedicated switch for it, so without that block every chat reply reasons before it answers — slower, and billed as output tokens. Autocomplete is unaffected either way, because Continue sends completions to DeepSeek's fill-in-the-middle endpoint, which only runs in non-thinking mode.

If you are still on the legacy config.json format, the equivalent shape uses title instead of name, a top-level tabAutocompleteModel object instead of the autocomplete role, and no roles array — but I recommend migrating to YAML, which is what Continue's documentation now ships as the default.

Step 3: Choose Your Model

DeepSeek's API now offers two tiers, and choosing the right one matters:

  • DeepSeek Flash (deepseek-flash, currently DeepSeek V4.1 Flash) — Fast, cheap, and the default for most tasks. Thinking mode is on by default, with low, high and max effort levels, and can be switched off per request for instant responses (the "Fast" entry above does exactly that). It also accepts images, so you can paste a screenshot of a broken layout straight into chat. At $0.15–0.30/M input and $0.60–1.20/M output (off-peak–peak), this is your daily-driver model.
  • DeepSeek V4 Pro (deepseek-v4-pro) — Higher-capacity reasoning model, text only (no image input). Use it for architectural design discussions, deep refactors, complex debugging, and longer-horizon code reviews. Pricier than Flash (currently $0.66–1.32/M input and $1.98–3.96/M output, off-peak–peak), but worth it for the hard problems. DeepSeek briefly announced that V4 Pro requests would be routed to V4.1 Flash from 14 September 2026, then reversed course after user pushback — it stays available at its own rates until a V4.1 Pro arrives.

I recommend using Flash for autocomplete and most chat interactions, and V4 Pro for when you genuinely need deeper reasoning. The configuration above wires this up directly — Continue's model picker lets you swap between them with a dropdown in the chat panel.

Method 2: DeepSeek via Ollama (Local, Free)

If you prefer not to send your code to any external server — or if you simply want zero ongoing cost — you can run DeepSeek Coder locally through Ollama. The trade-off is that you need decent hardware and the model quality is limited by what your machine can handle. Note that V4.1 Flash's open weights are not a laptop option: at 552B parameters it needs data-centre hardware, and the only Ollama library entry for it (deepseek-v4.1-flash:cloud) runs on Ollama's servers, not yours, so it does nothing for privacy. For genuinely local use, DeepSeek Coder V2 remains the practical pick.

Step 1: Pull the DeepSeek Coder Model

Open a terminal and run:

ollama pull deepseek-coder-v2:16b

This pulls the 16-billion parameter "Lite" variant of DeepSeek Coder V2 (a Mixture-of-Experts model with roughly 2.4B active parameters). It downloads at ~8.9GB in its default 4-bit quantisation and is the right balance of quality and performance for most developer machines — it runs comfortably on 8–12GB of VRAM, or on Apple Silicon with 16GB+ unified memory. The :latest tag currently points to the same 16B model.

If you have a workstation with serious GPU memory (think 80GB+ across one or more cards) and want the full-quality 236B model, pull instead:

ollama pull deepseek-coder-v2:236b

Note that there is no deepseek-coder-v2:lite tag in the Ollama library — the 16B model is the Lite variant, just published under its parameter count. Verify what you have by running ollama list; you should see deepseek-coder-v2:16b in the output.

Step 2: Configure Continue for Local DeepSeek

Update your config.yaml to point to your local Ollama instance:

models:
  - name: DeepSeek Coder V2 (Local)
    provider: ollama
    model: deepseek-coder-v2:16b
    roles:
      - chat
      - edit
      - apply
  - name: DeepSeek Coder V2 Autocomplete
    provider: ollama
    model: deepseek-coder-v2:16b
    roles:
      - autocomplete
    autocompleteOptions:
      debounceDelay: 500
      maxPromptTokens: 2048

Ensure Ollama is running — visit http://localhost:11434 in your browser to confirm you see "Ollama is running". Continue will automatically connect to the local Ollama API at that address and route all requests to your machine. No API key is required for local Ollama.

Step 3: Verify It Works

Open the Continue chat panel in VS Code and type a simple prompt like "Write a Python function that reverses a string." If you see a response from DeepSeek Coder, your local setup is working. Then open a code file and start typing — you should see tab-autocomplete suggestions appearing after a brief delay.

Configuring Autocomplete

Tab-autocomplete is the feature that makes the biggest difference in day-to-day coding. In Continue's config.yaml, autocomplete tuning lives under the autocompleteOptions block on whichever model has the autocomplete role (see the YAML examples above). Here is how to optimise DeepSeek's autocomplete behaviour:

  • Debounce delay. Set autocompleteOptions.debounceDelay: 300 on your autocomplete model. This waits 300 milliseconds after you stop typing before sending a completion request. Too low and you waste API calls (or GPU cycles) on partial words. Too high and suggestions feel sluggish. I find 300ms is the sweet spot for the DeepSeek API; bump it to 500ms if running locally on modest hardware.
  • Multiline completions. Set autocompleteOptions.multilineCompletions: "always" to let DeepSeek suggest entire function bodies, multi-line conditionals, and complete code blocks rather than just single lines. Flash and DeepSeek Coder V2 both produce well-structured multi-line completions that are often correct on the first suggestion.
  • Prompt token budget. Set autocompleteOptions.maxPromptTokens to control how much surrounding context Continue sends with each completion request. The default of 1024 is fine for the API. For local Ollama, keeping it at 1024–2048 keeps autocomplete responsive on modest GPUs; most inline completions do not need more context than that.
  • Fill-in-the-middle (FIM). DeepSeek Coder V2 and DeepSeek Flash both support FIM, meaning the model analyses code both before and after your cursor to generate contextually appropriate suggestions. For the API model, Continue's DeepSeek provider sends completions to the dedicated FIM endpoint (/beta/completions), which runs Flash in non-thinking mode, so autocomplete stays fast. This is particularly useful when you are writing code in the middle of an existing function — the suggestions will respect the surrounding code structure rather than treating your cursor position as the end of the file.

Chat-Based Workflows

Beyond autocomplete, the chat interface is where DeepSeek provides tremendous value. Here are the workflows I use most frequently with DeepSeek in VS Code:

Code Explanation

Select a block of unfamiliar code — perhaps something inherited from a colleague or pulled from a library — and press Ctrl+L to send it to Continue's chat. Ask: "Explain what this code does step by step. Highlight any potential issues." DeepSeek Flash in thinking mode (or V4 Pro for denser code) handles this exceptionally well. It correctly identifies design patterns, traces control flow, and flags common issues like missing error handling, race conditions, or inefficient algorithms. The explanations are clear and well-structured, often rivalling what you would get from GPT-class models.

Refactoring

Select a function and ask DeepSeek to refactor it. For example: "Refactor this function to use async/await instead of callbacks. Add TypeScript types. Extract the configuration into a separate object." DeepSeek Flash produces clean, idiomatic refactored code across JavaScript, TypeScript, Python, Go, and Rust. It occasionally over-engineers the solution — adding abstraction layers that are unnecessary for simpler functions — but a follow-up prompt like "simplify this, keep it under 30 lines" corrects that tendency quickly.

Test Generation

Paste a function and ask: "Write unit tests for this function using Jest. Cover the happy path, edge cases (empty input, null values, type mismatches), and error conditions." DeepSeek Flash generates solid test suites that cover the main execution paths. It handles standard testing frameworks well — Jest, Vitest, pytest, Go's testing package, Rust's built-in tests. Where it sometimes falls short is in generating truly creative edge cases; it covers the obvious ones but may miss domain-specific boundary conditions that a human tester would catch.

Debugging

Paste an error message along with the relevant code and ask: "What is causing this error and how do I fix it?" This is one of DeepSeek's strongest use cases. Common error patterns, stack traces, and framework-specific issues are well-represented in its training data. In my testing, DeepSeek correctly diagnosed the root cause on the first attempt roughly 80% of the time — comparable to Claude and OpenAI's frontier models for standard debugging tasks.

DeepSeek Flash vs V4 Pro: When to Use Which

DeepSeek has consolidated its API into two tiers. Understanding when each one earns its keep saves you both money and latency:

Aspect DeepSeek Flash (V4.1) DeepSeek V4 Pro
Primary strength Speed, cost, everyday tasks Deep reasoning, longer-horizon planning
Best for Autocomplete, quick chat, refactoring, test generation Architecture design, complex debugging, code review of large diffs
Modes Thinking (default; low/high/max effort) or non-thinking; image input Thinking (low/high/max effort); text only
Context window 1M tokens 1M tokens
Max output 384K tokens 384K tokens
Response speed Fast in non-thinking mode (typical 200–500ms first token) Slower; reasoning takes seconds, not milliseconds
Token cost (cache miss, off-peak–peak) $0.15–0.30/M input · $0.60–1.20/M output $0.66–1.32/M input · $1.98–3.96/M output

The practical recommendation: keep Flash as your default for autocomplete and most chat work, and reach for V4 Pro only when a problem genuinely needs deeper reasoning — refactors that touch many files, architectural trade-offs, or debugging that requires holding a lot of context in mind at once. The dual-model setup is already reflected in the configuration examples above.

Cost Breakdown: How $2/Month Actually Works

The $2/month figure is not marketing — it is based on real-world usage tracking. Here is how the maths works:

DeepSeek's API pricing below was last checked on 24 September 2026. Since 16 August 2026 DeepSeek bills by the hour: peak rates apply 01:00–04:00 and 06:00–10:00 UTC on weekdays (Chinese office hours), and everything else — evenings, weekends, Chinese public holidays — is off-peak at half price. Published rates change without much notice, so re-check the official pricing page before budgeting a new setup.

  • DeepSeek Flash (deepseek-flash): $0.30 per million input tokens (cache miss), $0.006 per million input tokens (cache hit), $1.20 per million output tokens at peak; $0.15, $0.003 and $0.60 off-peak.
  • DeepSeek V4 Pro (deepseek-v4-pro): $1.32 per million input tokens (cache miss), $0.044 per million input tokens (cache hit), $3.96 per million output tokens at peak; $0.66, $0.022 and $1.98 off-peak.

The cache-hit pricing is the genuinely interesting line. DeepSeek's API automatically caches repeated prompt prefixes — system prompts, file headers, code you've already discussed — and bills repeated tokens at 2–3% of the cache-miss rate. For tab-autocomplete and iterative chat, where the prompt prefix barely changes between requests, your effective bill drops dramatically below the headline numbers. The flip side is that output tokens — including the reasoning DeepSeek produces in thinking mode — are now the expensive line, which is why the "Fast" chat entry above switches thinking off for routine questions.

Compare this to the mid-tier frontier models: OpenAI's GPT-6 Sol and Anthropic's Claude Sonnet 5 both list $2 per million input tokens and $10 per million output tokens, and the flagship GPT-6 Astra is $10 and $50. Against those, DeepSeek Flash is roughly 7–13× cheaper on input and 8–17× cheaper on output depending on the hour. It is no longer the cheapest API by list price — OpenAI's budget-tier GPT-6 Luna is $0.10 and $0.50 — but DeepSeek's pitch is frontier-class coding quality at those prices, not the smallest model that happens to be cheap.

In a typical coding day, I make roughly 50–100 chat interactions and receive several hundred autocomplete suggestions. That works out to approximately 200,000–400,000 tokens per day in combined input and output, the large majority of it cached input. Most of the cost is the output: 50,000-odd generated tokens a day is $0.03–0.06 on its own. With caching and thinking off for routine chat, that comes to around $0.04 to $0.10 per day depending on how much of it lands in peak hours, or $1.50 to $3 per month. Leave thinking on for every chat and lean on V4 Pro and you can double that — still far short of a Copilot Pro seat.

If you use the local Ollama method, your cost is literally zero (beyond the electricity to run your GPU). The trade-off is that you need capable hardware and the local models are slightly less powerful than what the API provides.

For comparison, here is what the alternatives cost:

  • GitHub Copilot Pro: $10/month (Pro+ $39, Max $100 — every paid plan now bundles a monthly allowance of AI credits that meters premium model usage)
  • GitHub Copilot Business: $19/month per seat (Enterprise $39)
  • Claude: $20/month for the Pro subscription, or $2 input / $10 output per million tokens for Claude Sonnet 5 via the API
  • ChatGPT Plus (for coding): $20/month
  • DeepSeek via API: ~$2/month for typical usage
  • Gemma 4 via Ollama (local): Free

DeepSeek vs Copilot vs Claude vs Gemma 4 (Local)

This is the comparison that matters. Here is an honest assessment based on weeks of side-by-side testing:

Feature DeepSeek Flash (API) GitHub Copilot Claude Sonnet 5 (API) Gemma 4 (Local)
Monthly cost ~$2 $10 (Pro) ~$5-15 (API usage) Free
Autocomplete quality Very good Excellent Good (via Continue) Good
Autocomplete speed Fast (200-500ms) Very fast (100-300ms) Moderate (300-800ms) Hardware-dependent
Chat / reasoning quality Very good Very good Excellent Good
Code explanation Excellent Very good Excellent Good (27B model)
Test generation Good Good Excellent Adequate
Language breadth Strong in popular languages Broad across all languages Strong across all languages Strong in popular languages
Privacy Code sent to China-based servers Code sent to GitHub/Microsoft Code sent to Anthropic Fully local
Offline availability No No No Yes
Setup effort Moderate (API key + Continue) Minimal (install + sign in) Moderate (API key + Continue) High (Ollama + model + Continue)
Project context awareness Limited to active file + prompt Indexes workspace Limited to active file + prompt Limited to active file + prompt

The takeaway is clear: DeepSeek occupies a unique position as the best value-for-money AI coding assistant. It is not the absolute best in any single category, but it is remarkably close to the best in most categories at a tenth of the price. If you are a budget-conscious developer, a student, or someone who simply objects to paying $10/month when a $2/month option exists that covers 90% of your needs, DeepSeek is the obvious choice.

Limitations and Considerations

DeepSeek is impressive for its price, but it is not without drawbacks. You should be aware of these before committing:

API Reliability

DeepSeek's API has experienced periodic outages and rate limiting, particularly during peak hours. When DeepSeek went viral in early 2025, the API was frequently unavailable for hours at a time. Reliability has improved significantly since then, but it is still not as consistently available as OpenAI's or Anthropic's APIs. If you rely on AI coding assistance for production work, consider having a fallback (a local Ollama model, for instance).

Privacy and Data Concerns

This is the elephant in the room. DeepSeek is a Chinese company, and when you use the API, your code is transmitted to servers based in China. DeepSeek's privacy policy states that data may be stored and processed in the People's Republic of China. For personal projects, open-source work, and learning, this is unlikely to be a practical concern. For proprietary corporate code, code subject to regulatory compliance (GDPR, HIPAA, SOX), or sensitive intellectual property, this is a serious consideration that may disqualify DeepSeek entirely.

If privacy is a hard requirement but you still want DeepSeek's capabilities, use the Ollama local method. Running DeepSeek Coder locally means no data leaves your machine — ever. The model quality is somewhat lower than the full API models, but it eliminates the data residency concern completely.

Content Filtering and Censorship

DeepSeek's models include content filters that occasionally interfere with legitimate coding tasks. In my testing, this was rare for standard development work, but it can surface when working with security-related code (penetration testing tools, encryption implementations), content that touches on politically sensitive topics (relevant for some NLP and content moderation projects), or certain medical or legal domain code. If you hit a refusal, rephrasing the prompt usually resolves it.

Language and Framework Coverage

DeepSeek excels at Python, JavaScript/TypeScript, Go, Rust, Java, and C++. It handles these languages at a level very close to Copilot. Where it falls behind is in less common languages (Elixir, Haskell, OCaml, Kotlin Multiplatform) and in very new frameworks or libraries that were released after its training cutoff. Copilot's advantage here comes from its continuous learning pipeline and GitHub's vast code corpus.

Frequently Asked Questions

Is DeepSeek safe to use for coding?

For personal projects, open-source contributions, and learning — yes, it is perfectly safe. The models produce high-quality code and the API functions reliably for most use cases. The primary concern is data privacy: your code is sent to DeepSeek's servers in China when using the API. If you work with proprietary or regulated code, either use the local Ollama method (which keeps everything on your machine) or check with your organisation's security team before using the cloud API. From a code quality perspective, DeepSeek's suggestions are comparable to other leading AI models — always review generated code before committing it, regardless of which AI tool produced it.

How does DeepSeek's $2/month compare to GitHub Copilot's $10/month in practice?

The day-to-day experience is surprisingly close. For autocomplete, Copilot is still faster and slightly more contextually aware (it indexes your entire workspace), but DeepSeek Flash produces relevant completions for most standard coding tasks. For chat-based workflows — explaining code, refactoring, debugging, writing tests — the quality gap is negligible with Flash, and V4 Pro closes it further on harder reasoning tasks. Where Copilot clearly wins is in setup simplicity (one-click install versus API key configuration), reliability (near-100% uptime versus occasional DeepSeek outages), and workspace-wide context awareness. Whether the $8/month difference justifies those advantages depends entirely on your priorities and budget.

Can I switch between DeepSeek and other models in the same VS Code setup?

Yes. The Continue extension supports multiple model providers simultaneously. You can configure DeepSeek, Claude, OpenAI, and local Ollama models all in the same config.yaml file and switch between them with a dropdown in the Continue panel. This is actually the setup I recommend: use DeepSeek as your primary model for cost efficiency, keep a local Ollama model as a fallback for when the API is unavailable, and optionally add a Claude or OpenAI model for tasks where you want the absolute best quality regardless of cost.

Related Tutorials

Sources