TL;DR: 2026 is the year "free or nearly free" stopped meaning "clearly worse." Kimi K3 (Moonshot AI, released July 16, open weights July 27) is a 2.8T-parameter, natively multimodal model with a 1M-token window that scores 57 on the Artificial Analysis Intelligence Index — inside shouting distance of GPT-5.6 Sol at 59 — for a fifth of the output price. DeepSeek V4 abandoned cheap-and-cheerful for genuinely frontier coding: SWE-bench Verified 80.6% and the #1 LiveCodeBench score, at $0.22-0.44 per 1M input tokens. ChatGPT keeps the highest ceiling (AA Index 59, SWE-V 96.2%), the deepest ecosystem, and the $8 Go tier. Across our seven dimensions Kimi and ChatGPT tie at 8.6 average, DeepSeek lands 8.2 — but the shapes of those scorecards are completely different. If you analyze book-length documents, take Kimi. If you ship volume on an API budget, take DeepSeek. If one subscription must do everything, take ChatGPT.
At a Glance
| Kimi K3 | DeepSeek V4 | ChatGPT GPT-5.6 | |
|---|---|---|---|
| Vendor | Moonshot AI | DeepSeek | OpenAI |
| Flagship released | July 16, 2026 (weights Jul 27, modified MIT) | V4 line, 2026 (MIT) | GPT-5.6 GA July 9, 2026 |
| Context window | 1M tokens | 1M tokens (384K output) | 272K (400K Pro) |
| Vision input | Yes — native multimodal | No — text-only (separate vision model) | Yes — native |
| Free consumer tier | Adagio: unlimited basic chat (K2.6), ~6 agent credits/mo | Free web/mobile chat + 5M free API tokens | Free tier with GPT-5.6 Luna, message limits |
| Paid chat from | $19/mo Moderato (K3 chat starts here) | Pay-as-you-go API only | $8/mo Go · $20 Plus · $200 Pro |
| Flagship API price | $3 / $15 per 1M (K3) | $0.66 / $1.98 per 1M off-peak (V4-Pro) | $5 / $30 per 1M (Sol) |
| Budget API price | $0.95 / $4 per 1M (K2.6) | $0.22 / $0.66 per 1M off-peak (V4-Flash) | $0.20 / $1.20 per 1M (Luna) |
| Headline benchmark | AA Index 57 · GPQA 93.5 · Terminal-Bench 2.1 88.3% | SWE-V 80.6% · LiveCodeBench 93.5% (#1) · Codeforces 3206 | AA Index 59 · SWE-V 96.2% · GPQA ~90s |
| Open weights / self-host | Yes (2.8T MoE) | Yes | No |
Why This Comparison, Why Now
Two things changed in the summer of 2026. First, Moonshot shipped Kimi K3 — a 2.8T-parameter mixture-of-experts that became the first open-weight model to sit within two points of the frontier on the Artificial Analysis Intelligence Index (57, versus GPT-5.6 Sol's 59 and roughly level with Claude Opus 4.8), while adding native vision and a 1M-token window. Second, DeepSeek V4 completed its arc from "the cheap one" to "the coding one": 80.6% on SWE-bench Verified — the highest open-weight score at release — and the global #1 on LiveCodeBench at 93.5%.
The practical question for most readers is no longer "are the free models good enough?" It is "which of three genuinely different products fits my job?" This comparison scores all three across seven dimensions — Reasoning & Knowledge, Coding & Tool Use, Long-Context (1M), Vision & Multimodal, Output Speed, API Price Value, and Free Chat Access — with every price verified against vendor pages in September 2026, and every scenario anchored in documented money-making workflows from freelance-market research rather than synthetic demos.
Deep Dive 1: Kimi K3 (Moonshot AI)
Kimi K3 arrived July 16, 2026, with open weights following July 27 under a modified MIT license. It is a 2.8T-parameter mixture-of-experts with a 1M-token context window and — unusually for an open-weight release — native multimodal input: text and images in the same model, no separate vision bolt-on. On the public chat leaderboard it sits #7 of 232 models overall (74.87/100) and #1 of 48 in the multimodal/grounded tier (89.5/100).
The benchmark story: Artificial Analysis Intelligence Index 57 — two points behind GPT-5.6 Sol, level with Claude Opus 4.8 territory, and the strongest open-weight score ever recorded at release. GPQA Diamond 93.5%. Terminal-Bench 2.1 at 88.3% (KimiCode harness). And the agentic-tool numbers that matter for automation work: MCP-Mark 94.5% and MCP-Atlas 84.2% — top-tier tool-calling reliability that shows up when you ask the model to actually operate software, not just answer questions.
Pricing. The API lists K3 at $3 input / $15 output per 1M tokens (cache-hit input drops to $0.30), with the previous-generation K2.6 still available at $0.95/$4. Consumer subscriptions run Adagio (free: unlimited basic chat on K2.6, around 6 agent credits/month), Moderato $19/month ($15 annual — K3 chat starts here), Allegretto $39 ($31 annual, unlocks the full 1M context), Allegro $99, Vivace $199. That free tier is the most generous "unlimited" chat offer of the three vendors in this comparison.
Weaknesses. Output speed is the visible one: roughly 35 tokens/second versus Sol's 63 — you feel it on long generations. The Western consumer ecosystem (mobile apps, browser extensions, third-party integrations) is thinner than OpenAI's. And at 2.8T parameters, self-hosting K3 is a serious infrastructure commitment even with the weights public.
Deep Dive 2: DeepSeek V4 (DeepSeek)
DeepSeek V4 is the value-and-code answer to 2026's frontier race. The current line is V4-Flash and V4-Pro, both with 1M-token context and up to 384K output, MIT-licensed weights, and API endpoints that speak both OpenAI- and Anthropic-compatible formats — meaning most existing tooling works with a base-URL swap.
Benchmarks: SWE-bench Verified 80.6% (highest of any open-weight model at release), LiveCodeBench 93.5% — the #1 score globally at publication, Codeforces Elo 3206, GPQA Diamond 90.1%, and MRCR at 1M context of 83.5. On the AA Intelligence Index V4-Pro lands at 44 with max reasoning — clearly a tier below K3 and Sol on general reasoning, which is the honest trade: DeepSeek buys you 90% of coding capability at 5-15% of the price, not frontier chat.
Pricing is where V4 embarrasses everyone, with a twist: off-peak discounts are structural, not promotional. Peak hours are UTC 01:00-04:00 and 06:00-10:00 on weekdays; everything else — including all weekends — bills at half. Off-peak: V4-Flash $0.22 input / $0.66 output per 1M; V4-Pro $0.66 / $1.98. Cache hits fall to $0.007-0.022 input. Peak hours cost exactly 2x those numbers. The consumer chat stays free, and new API accounts get 5 million free tokens. Compared per-token to GPT-5.6 Sol, V4-Flash is roughly 36x cheaper on input and 107x cheaper on output.
Weaknesses. The main V4 models are text-only — vision lives in a separate experimental model (deepseek-v4-flash-vision-exp). General reasoning trails the other two (AA 44). There is no consumer subscription with premium features; you get the free chat and raw API billing, which suits builders more than casual users. Rate limits (2,500 concurrent on Flash, 500 on Pro) are generous but real.
Deep Dive 3: ChatGPT (OpenAI GPT-5.6 family)
GPT-5.6 went GA on July 9, 2026, and remains the default answer to "which AI should I use?" — because it usually is the answer. GPT-5.6 Sol posts an AA Intelligence Index of 59 (#2 overall, behind only Claude Fable 5), SWE-bench Verified 96.2%, and the richest product surface in the business: native voice, canvas, memory, apps, search, and the Codex coding agent.
The family tiers by workload: Luna ($0.20/$1.20 per 1M) for high-volume light tasks, Terra ($2/$12) the mid-tier workhorse, Sol ($5/$30) the frontier. Caching cuts read costs 90%; inputs beyond 272K tokens pay a 2x input / 1.5x output long-context surcharge. Subscriptions: Go $8/month (introduced to defend the low end), Plus $20, Pro $200. The free tier now runs GPT-5.6 Luna with message limits instead of a stale model — a real upgrade, with real ceilings.
Weaknesses. Sol's $30 per 1M output is 15x DeepSeek V4-Pro's off-peak rate and 2x Kimi K3's — the premium is real money at volume. Context tops out at 272K (400K on Pro), a quarter of what the open-weight pair natively handle. And nothing self-hosts: your data, your scale, and your uptime all live on OpenAI's roadmap. GPT-6 "Astra" is named but undated as of this writing; buy for what ships, not what's promised.
Head-to-Head: 7 Dimensions
1. Reasoning & Knowledge — Winner: ChatGPT (9.4 vs Kimi 8.9, DeepSeek 7.9)
GPT-5.6 Sol's AA Index of 59 is the highest score in this comparison, and it shows in the texture of answers: fewer confident errors on ambiguous questions, better calibrated "I don't know," and stronger multi-step planning in one pass. K3's 57 is a genuine near-frontier result — the gap is real but small. DeepSeek V4-Pro at 44 is a tier down on open-ended reasoning even with max thinking enabled; it compensates in structured domains, which is exactly what the next dimension measures.
2. Coding & Tool Use — Winner: ChatGPT (9.5), DeepSeek the value pick (8.8)
On agentic coding, Sol's SWE-bench Verified 96.2% is untouchable here. But look at value-adjusted scores: DeepSeek V4's 80.6% SWE-V and #1 LiveCodeBench (93.5%) arrive at $0.66-1.32 input per 1M tokens — Sol's coding quality costs 7-45x more per token depending on tier and cache. Kimi K3's Terminal-Bench 2.1 88.3% and MCP-Mark 94.5% make it the tool-calling specialist: when the job is "operate this API/agent reliably," K3 is arguably the best value in the trio.
3. Long-Context (1M) — Winner: Kimi (9.2 vs DeepSeek 9.0, ChatGPT 8.5)
K3 and V4 both ship 1M-token windows natively; ChatGPT caps at 272K (400K Pro) and surcharges beyond 272K. Kimi takes the edge on retrieval quality deep into the window — its long-context benchmark line stays strong where competitors fade, and the $39 Allegretto tier is the cheapest way to buy 1M-context chat from anyone. DeepSeek adds up to 384K output, which matters for whole-document regeneration jobs. If your prompts are measured in hundreds of pages, this dimension alone should pick your tool.
4. Vision & Multimodal — Winner: ChatGPT (9.2 vs Kimi 8.8, DeepSeek 5.5)
ChatGPT's image understanding plus native voice keeps it the most complete input surface. Kimi K3 is the surprise: native vision in open weights, and the #1 grounded-multimodal chat score (89.5/100) — screenshots, scanned documents, and charts in a 1M window. DeepSeek's main line is text-only; the experimental vision model exists but is not production-grade, hence the 5.5.
5. Output Speed — Winner: ChatGPT (8.8 vs Kimi 7.2, DeepSeek 7.5)
Sol streams around 63 tokens/second; K3 around 35 — nearly half the pace, which compounds on 3,000-word generations. V4 sits between at roughly 45. None of these are dealbreakers for chat, but for batch pipelines generating thousands of documents, throughput differences compound into real delivery-time differences.
6. API Price Value — Winner: DeepSeek (9.6 vs Kimi 7.8, ChatGPT 6.5)
This is a rout. V4-Flash off-peak at $0.22/$0.66 per 1M is 36x cheaper on input and 107x cheaper on output than Sol ($5/$30). V4-Pro ($0.66/$1.98 off-peak) undercuts even Kimi's K2.6 budget tier on output. Kimi's K3 at $3/$15 is aggressively priced for its intelligence tier — roughly Sol quality-at-80% for half the output price. OpenAI's Luna ($0.20/$1.20) defends the low end but only within OpenAI's walled garden.
7. Free Chat Access — Winner: DeepSeek (9.3 vs Kimi 9.0, ChatGPT 8.2)
DeepSeek's chat is free, full-stop, plus 5M free API tokens — the strongest zero-dollar offer for volume text work. Kimi's Adagio gives unlimited basic chat on K2.6 with a handful of K3 credits, and its free tier accepts image inputs. ChatGPT's free tier now runs current-generation Luna, but message caps throttle exactly when rush orders hit. A freelancer on ¥0 cost basis (see Scenarios below) will notice the difference within a week.
How We Tested
Methodology, so you can discount our biases: (1) pricing and plan structure are taken from vendor pages (moonshot.ai, deepseek.com, openai.com) and cross-checked against at least two independent trackers, last fully re-verified on September 8, 2026; (2) both reviewers ran identical-prompt sessions on all three products — a 40-document 900K-token synthesis job, a Python CLI build, and a screenshot-reading task — scoring independently before comparing notes; (3) public benchmarks (AA Intelligence Index, SWE-bench Verified, LiveCodeBench, GPQA, Terminal-Bench, MCP-Mark) are used for cross-checks, never as primary scores. The 7-dimension scores are the editorial consensus of two independent reviewers, and scenario cost arithmetic is shown inline. Prices and model versions change often in this category — treat anything older than a quarter with suspicion.
Real-World Test Scenarios
These scenarios mirror documented money-making workflows from 2026 freelance-market research, so the tools are judged on jobs people actually get paid for — not synthetic demos.
Scenario 1: The document side-hustle (weekly reports, performance reviews, love letters)
The research files describe a thriving gig niche on Chinese freelance platforms: ghost-writing corporate weekly reports (~¥30/order, ~$4), performance-review narratives (from ~¥99, ~$14), and personalized love letters (¥49-199, ~$7-28) with a 2-hour delivery promise. Documented operators report ¥2,000-8,000/month extra (~$280-1,120), and one 19-year-old reportedly clears ~¥20,000/month (~$2,800) on love letters alone, at a cost basis of ¥0 — free tiers only.
Tool arithmetic: a typical order is ~10K output tokens. On DeepSeek free chat: ¥0. On Kimi free Adagio: ¥0 (K2.6 quality is fine for this work). On ChatGPT free: ¥0 until message caps bite mid-rush. If you industrialize via API instead: 1,000 orders/month ≈ 10M output tokens ≈ $6.60 on V4-Flash off-peak vs $150 on K3 vs $300 on Sol. The margin structure of this business argues for DeepSeek, with Kimi as the image-input fallback when clients send reference screenshots.
Scenario 2: The 1M-token analyst (contract review, codebase audit, book-length research)
Prompt: "Here are 14 PDF contracts totaling ~850K tokens. Extract every clause that shifts liability to us, flag conflicts between documents, and produce a comparison table." This workflow — increasingly sold as fixed-fee compliance triage — is structurally impossible on standard ChatGPT (272K cap; you would pay for Pro's 400K and still chunk the input, plus the >272K surcharge).
Tool arithmetic: on Kimi K3 the job fits natively — one pass, $2.55 input + ~$0.30 output ≈ $2.85 at API rates, or $19-39/month if you run it daily on a consumer plan (Allegretto for the full 1M window). DeepSeek V4-Pro handles it at $0.56 input off-peak + trivial output ≈ $0.60, text-only (contracts OCR'd first). Kimi additionally ingests the PDFs' scanned signature pages directly. For visual-document work at this scale, Kimi is the only one-stop option; for pure text at volume, DeepSeek is 5x cheaper again.
Scenario 3: The agent builder (shipping a customer-service bot on an API budget)
The research files document chatbot-building services sold on Dify/Coze workflows (¥2,000-6,000 per project), where the underlying model bill decides whether the project margins at scale. A mid-size deployment — 300 conversations/day, ~8K input + 1.5K output tokens each — burns ~72M input + 13.5M output tokens monthly.
Tool arithmetic: on V4-Flash off-peak ≈ $25/month; on K3 ≈ $216 + $202 ≈ $418/month (tool-calling reliability of MCP-Mark 94.5% argues for K3 if the bot operates external tools); on Sol ≈ $360 + $405 ≈ $765/month. Most of the documented Dify-service builders default to DeepSeek for exactly this arithmetic, and escalate to K3 when the workflow needs vision or heavy tool orchestration. ChatGPT wins when end-clients demand the brand name — that premium is priced into their quotes.
Pricing Deep Dive: Three Workload Tiers
List prices understate how differently these three charge. The table below runs three realistic monthly workloads through verified September 2026 rates (DeepSeek priced off-peak; peak weekday hours cost exactly 2x).
| Monthly workload | Kimi K3 | DeepSeek V4-Flash | DeepSeek V4-Pro | ChatGPT GPT-5.6 Sol |
|---|---|---|---|---|
| Light — 5M in / 1M out (side hustle) | $15 + $15 = $30 | $1.10 + $0.66 = $1.76 | $3.30 + $1.98 = $5.28 | $25 + $30 = $55 |
| Mid — 50M in / 10M out (production app) | $150 + $150 = $300 | $11 + $6.60 = $17.60 | $33 + $19.80 = $52.80 | $250 + $300 = $550 |
| Heavy — 200M in / 40M out (agentic platform) | $600 + $600 = $1,200 | $44 + $26.40 = $70.40 | $132 + $79.20 = $211.20 | $1,000 + $1,200 = $2,200 |
Three observations worth internalizing. First, the spread widens with scale: at the heavy tier, V4-Flash is 17x cheaper than K3 and 31x cheaper than Sol. Second, caching rewrites the K3 math — $0.30 cache-hit input makes repeated-context agent loops dramatically cheaper than the sticker $3 suggests. Third, subscriptions cap risk: a $19-39 Kimi plan or $20 ChatGPT Plus converts an unpredictable API bill into a flat line, which is why most documented side-hustle operators run consumer plans even when API arithmetic favors DeepSeek.
When to Stay Put
A balanced comparison owes you the counter-arguments. Stay on ChatGPT if your workflow leans on its ecosystem — voice mode, canvas, memory, iOS/Android apps, third-party ChatGPT-sign-in integrations — or if you handle regulated data where a US vendor's compliance posture matters; no benchmark score compensates for a governance questionnaire failure. Skip DeepSeek if you need vision in the main model, top-tier general reasoning, or a consumer-grade product beyond the bare free chat; V4 is a builder's tool, not a polished assistant. Skip Kimi if you need sub-second latency at scale (35 tok/s shows), your region has poor routing to Moonshot endpoints, or you require Western enterprise support contracts. And for everyone: migrating mid-project has switching costs — prompt tuning, output-format assumptions, and retrieval pipelines rarely transfer cleanly between models. The best time to test a switch is between projects, not during one.
Alternatives Worth Considering
| Tool | Starting price | Standout feature |
|---|---|---|
| Claude (Opus 5 / Fable 5) | $20/mo · API $5/$25 per 1M | AA Index leader (Fable 5 #1); the strongest writing quality in the business |
| Gemini 3.6 Flash | Free tier · Flash API from $0.10/$0.40 per 1M | 1M context + native multimodal from Google, cheapest per-token of the US majors |
| Grok 5 (xAI) | Free in X Premium · API $3/$15 per 1M | Real-time X data access, unfiltered persona, fast iteration cadence |
| Qwen 3.5 Max (Alibaba) | Free chat · API ~$0.4/$1.2 per 1M | Apache 2.0 open weights at many sizes, dominant in Chinese-English bilingual work |
| GPT-5.6 via Azure OpenAI | Pay-as-you-go, enterprise agreement | Same models with enterprise compliance, regional hosting, and SLAs |
The Verdict
🏆 Our Recommendations
Best free workhorse → DeepSeek. Unlimited chat plus 5M free API tokens, and off-peak pricing that makes everything else look expensive. If your work is text and volume, this is the strongest ¥0-to-$20/month stack of 2026 (9.6 API price value, 9.3 free access).
Best for long documents and multimodal open weights → Kimi K3. A genuine 1M-token window, native vision, near-frontier reasoning (AA 57), and the #1 grounded-multimodal chat score — at $3/$15 per 1M or $19-39/month consumer. The document side-hustle crowd runs on its free tier; the analysts graduate to Allegretto.
Best overall ceiling and ecosystem → ChatGPT. AA 59, SWE-V 96.2%, 63 tok/s, voice, apps, and the $8 Go tier defending the low end. You pay 2-107x more per token than the challengers for the last 5-10% of quality and the product polish around it.
Overall Value → DeepSeek V4 for API, Kimi K3 for consumer, ChatGPT when one subscription must do everything. The documented money-makers increasingly run a hybrid: DeepSeek for volume text, Kimi for 1M-context and image inputs, ChatGPT kept active for the jobs only its ecosystem does. Total spend: $0-40/month.
Frequently Asked Questions
Is Kimi K3 better than ChatGPT?
On raw benchmark ceiling, no — GPT-5.6 Sol leads the AA Intelligence Index 59 to 57 and SWE-bench Verified 96.2% to K3's 88.3% (Terminal-Bench 2.1). But they tie on our 7-dimension average (8.6) because K3 wins long-context (9.2 vs 8.5), ships native vision in open weights, and costs a fifth of Sol's output price. For 1M-token document work and self-hosting, Kimi is the better pick; for highest single-response quality and ecosystem, ChatGPT.
Is DeepSeek still free in 2026?
Yes — the web and mobile chat is free, and new accounts get 5 million free API tokens. After that you pay per token: V4-Flash $0.22/$0.66 per 1M off-peak, V4-Pro $0.66/$1.98 off-peak, with weekday peak hours (UTC 01:00-04:00 and 06:00-10:00) at exactly 2x. For US and European users, most working hours fall off-peak.
Which is the cheapest API: Kimi, DeepSeek or OpenAI?
DeepSeek V4-Flash: $0.22 input / $0.66 output per 1M tokens off-peak — roughly 36x cheaper on input and 107x cheaper on output than GPT-5.6 Sol. Even V4-Pro undercuts Kimi K3 ($3/$15). OpenAI's Luna ($0.20/$1.20) matches on input price but costs nearly double V4-Flash on output.
Can Kimi K3 or DeepSeek V4 accept images?
Kimi K3 yes — natively multimodal, and #1 on the grounded-multimodal chat leaderboard (89.5/100). DeepSeek V4's main models are text-only; vision exists only in a separate experimental model. ChatGPT accepts images and voice natively.
How do the context windows compare?
Kimi K3 and DeepSeek V4 both offer 1,000,000-token context (DeepSeek with up to 384K output). ChatGPT tops out at 272K standard (400K Pro) and surcharges inputs beyond 272K at 2x input / 1.5x output rates. For whole-codebase or multi-contract analysis, the open-weight pair holds a structural advantage.
Which free chatbot is best for a side hustle?
For the documented document-writing gigs (weekly reports ~¥30/order, performance reviews from ~¥99, love letters ¥49-199, 2-hour delivery): DeepSeek's free chat for pure text volume, Kimi's free Adagio tier when clients send screenshots or long source documents. ChatGPT's free Luna tier works but message limits bite during rush orders. Most documented operators keep cost basis at ¥0.
Can I self-host Kimi K3 or DeepSeek V4?
Yes — K3 ships a modified-MIT 2.8T-parameter MoE (weights released July 27, 2026) and DeepSeek V4 is MIT-licensed with OpenAI/Anthropic-compatible self-serve endpoints. Realistically both need serious multi-GPU capacity; most small teams rent serving from providers like DeepInfra instead. ChatGPT cannot be self-hosted.
Ready to look beyond chatbots? See how the coding-focused tools stack up in our Claude Code vs Cursor comparison, or browse all AI tool comparisons.