DeepSeek
Quick facts
- Operator
- DeepSeek
- Docs
- https://api-docs.deepseek.com/
- Engine class
- Retrieval-augmented chat with a user-controlled web search mode on DeepSeek's official consumer surfaces.
- Current consumer models
- DeepSeek V4 Pro and V4 Flash are available through the official app and web product as of August 2026.
- Crawler disclosure
- DeepSeek had not publicly documented a crawler or user-agent name as of August 29, 2026.
- Programmatic search
- The Responses API supports a server-side web_search tool and reports web_search_call actions.
- Evidence limit
- DeepSeek does not publish its consumer search index, source-ranking formula, or whether API and consumer results are equivalent; current measurements are observational snapshots.
1. What DeepSeek is
DeepSeek is a conversational AI product with an optional web search mode. When search is enabled, the hosted product can retrieve current web material, combine it with a model-generated answer, and display source information. This makes the product a retrieval-augmented generative engine within generative engine optimization, rather than a conventional search-results page with ranked links.
The name also refers to a model family, an API, and open weights. These surfaces share technology and branding, but they do not share a proven source index or ranking system. DeepSeek’s V4 announcement made V4 Pro and V4 Flash available through the official web product, app, API, and open-weight distribution. It did not say that self-hosted or third-party deployments inherit the official product’s search system.
| Surface | Who controls retrieval | What a citation test measures |
|---|---|---|
| Official DeepSeek web and app | DeepSeek controls the hosted search and answer experience | Visibility in the consumer product at a recorded date, mode, language, and region |
| Official DeepSeek API | The request selects an available tool and model; DeepSeek runs documented server tools | Behavior of that API configuration, not consumer ranking parity |
| Open-weight or self-hosted model | The deployer supplies the index, search provider, fetcher, prompts, and interface | The deployer’s retrieval system |
| Third-party product using DeepSeek | The third party combines a DeepSeek model with its own data and tools | The third party’s product behavior |
This distinction is especially important when comparing DeepSeek with other ecosystems that combine models and products, such as Tongyi and Doubao. A model benchmark can describe reasoning or tool use, but it cannot reveal which web pages the official consumer product retrieves or credits.
2. How DeepSeek web search works
DeepSeek launched Internet Search on its website on December 10, 2024. The launch notice said the product could extract several keywords and search them in parallel. The app followed in January 2025, adding web search, DeepThink, file upload, and synchronized chat history.
That public description supports a broad answer loop: interpret the question, create search queries, retrieve candidate pages, select evidence, synthesize an answer, and present sources. DeepSeek’s current privacy policy says it uses third-party APIs to provide search. It also says DeepSeek shares users’ input keywords with those providers. The policy does not name the provider or disclose the ranking model, reranking features, or rules that decide which retrieved page receives visible credit.
| Stage | Public evidence | What remains unknown |
|---|---|---|
| Search activation | The consumer product exposes a web search mode | Whether every enabled turn invokes retrieval |
| Query generation | The 2024 launch described several extracted keywords and parallel searches | Current V4 query count, rewriting rules, and language routing |
| Candidate retrieval | DeepSeek says third-party APIs provide search; independent tests observe current web information and sources | Provider identity, corpus boundaries, and candidate ranking |
| Evidence selection | Answers use some retrieved material and omit other candidates | Reranking weights and source-quality thresholds |
| Attribution | Consumer tests record source links or source labels | The rule that converts retrieved evidence into a visible citation |
Search mode, model mode, and conversation state should be recorded separately. A May 2026 consumer-interface study enabled both Deep Think and Web Search for DeepSeek V4 Flash, but that setup does not prove that every combination behaves the same. File uploads and prior turns can also place material in the context without open-web discovery.
Language is another independent variable. Chinese and English prompts may generate different queries, retrieve different domains, and use different entity names. A multilingual GEO test should keep the intent constant while recording query language, source language, location, and the exact product surface. Comparisons with Baidu AI Search, Quark AI Search, and Kimi require the same controls.
3. Crawlers and user-agents
DeepSeek has not publicly named a crawler or user-agent for consumer search in the official materials reviewed on August 29, 2026. The Tow Center’s March 2025 audit also reported that DeepSeek’s crawler was not publicly known. Stanford’s December 2025 transparency assessment reached the same conclusion after examining DeepSeek’s disclosures.
The absence of a disclosed identity does not mean that access never occurs. Search systems can discover a page through background indexing, a user-triggered fetch, a third-party search provider, a syndicated copy, or another public page that quotes the original. Each path creates a different access-control problem.
| Access role | Public DeepSeek identifier | Publisher control that can be verified today |
|---|---|---|
| Model-training collection | None documented | Apply a general data-use policy and review logs; do not guess a DeepSeek-specific UA |
| Search indexing | None documented | Maintain ordinary search accessibility and inspect verified requests |
| User-triggered page fetch | None documented | Test the public product with controlled pages and compare server logs |
| Third-party search or syndication | No provider disclosed | Audit copies, canonical links, and the source that the answer actually credits |
For an access audit, capture the complete user-agent string, source IP, requested URL, response status, rendered content, timestamp, and request pattern. Verify IP ownership or published ranges when available. A string containing DeepSeek is not enough, because user-agent text is easy to spoof. The general AI crawler audit method still applies even when the vendor has not supplied a named bot.
An invented robots.txt directive can create false confidence while leaving the real retrieval path untouched. Baidu AI Search should be evaluated separately because a documented search crawler or existing index relationship on one Chinese platform does not establish DeepSeek’s path.
4. Citation preferences
DeepSeek does not publish a consumer source-ranking formula. Its official product announcements confirm web search, but they do not state universal preferences for first-party sites, recent pages, Chinese domains, scholarly sources, or a particular content format. Those hypotheses need query-level evidence.
Visible outcomes should be classified before any preference is inferred:
| Outcome | Observable evidence | What it means |
|---|---|---|
| Citation or source link | A page or domain is linked as support for an answer | The interface gave that source visible credit |
| Named source without a link | A publisher, author, or brand is named but no page is linked | A mention occurred; the underlying page may be unclear |
| Uncredited answer | No source is shown | The answer may rely on model memory, retrieval, supplied context, or a mixture |
| Misattribution | The answer links or names the wrong origin for supplied material | Visible credit is inaccurate even if the prose is plausible |
The difference between citation, mention, and link matters because a brand can appear without its page being retrieved, and a retrieved page can influence an answer without receiving a visible link. Search presence, source selection, and attribution should therefore be measured as separate events.
The clearest DeepSeek-specific accuracy evidence comes from a failure audit. In 2025, the Tow Center supplied DeepSeek Search with excerpts from 200 news articles and asked for the original headline, publisher, date, and URL. DeepSeek misattributed the source in 115 cases. The result shows that visible source credit can be wrong. It does not identify a ranking factor, and it predates the current V4 product.
A broader 2026 Chinese-language generative search study observed an average of 9.2 citations per DeepSeek answer on both the web and app surfaces. Among 580 paired queries, the web and app results had 40.8% exact-URL overlap and 50.7% domain overlap. This is evidence of surface-level variation, not a disclosed preference or a stable platform-wide rate. The study is a preprint based on a June and July 2026 snapshot, and its public dataset should be used with those boundaries intact.
Practical experiments can test source type, freshness, language, and passage structure, but they should not turn correlation into a platform rule. Direct, self-contained evidence remains easier to inspect and quote after retrieval, which is the editorial logic behind writing for AI citation. It is not a disclosed DeepSeek ranking weight.
5. API and integration
DeepSeek’s API now has two different tool paths. Ordinary function calling lets the model request a developer-defined function; the developer then executes it and returns the result. The current Responses API also supports a server-side web_search tool executed by DeepSeek.
A minimal tool declaration looks like this:
{
"tools": [{ "type": "web_search" }],
"tool_choice": { "type": "web_search" }
}
The documented response can include web_search_call output items. Their actions can describe searching, opening a page, or finding text within a page. The current reference does not define a separate citation or source-URL object contract, so a collector should not assume every search action yields a machine-readable citation. The API is stateless, so a multi-turn test must send the relevant history again. Model support and fields can change; the request, returned model, tool choice, search actions, answer text, and any displayed or returned URLs should be retained together.
| Surface | Search capability | Measurement use | Main limitation |
|---|---|---|---|
| Consumer web or app | User-facing Web Search | Measures the actual hosted experience | Automation and result extraction may be limited |
| Responses API | DeepSeek-hosted web_search server tool | Repeated prompts can record whether search was called and what the answer returned | No documented consumer-ranking equivalence or dedicated citation-object contract |
| Chat Completions with function calling | Developer supplies the search function and results | Tests a controlled retrieval pipeline | Measures the developer’s search stack |
| Open-weight deployment | Deployer chooses every retrieval component | Supports reproducible custom experiments | Does not reproduce official DeepSeek Search |
The API became materially closer to a search-measurement surface in 2026, but it remains a proxy for the consumer product. AI citation tracking should store consumer and API observations in separate series. A model name, location, tool configuration, or account context can change the candidate set before citation selection begins.
6. History and timeline
DeepSeek’s GEO-relevant history is a sequence of product and tool changes rather than a simple model-release list.
| Date | Change | Effect on web visibility |
|---|---|---|
| July 25, 2024 | The API added function calling | Developers could connect DeepSeek models to an external crawler or search function, but the tool remained developer-supplied |
| December 10, 2024 | Internet Search launched on the official website | The consumer product gained live retrieval; the announcement explicitly said the API did not yet support search |
| January 15, 2025 | The official mobile app launched with web search and DeepThink | Search became available across the hosted app and web experience |
| September 22, 2025 | DeepSeek-V3.1-Terminus reported that Search Agent performance had improved | Search-agent behavior changed with a model update, although the announcement did not disclose ranking details |
| April 24, 2026 | V4 Preview reached the app, web, API, and open-weight channels | Surface and model labels needed to be recorded separately as the same model family spread across different retrieval systems |
| August 13, 2026 | V4 Pro GA reached app, web, and API; native Responses API support launched | The official API gained a documented server-side web-search path suitable for separately labeled tests |
The API change log is the best current record for model and interface changes. Search experiments should preserve the collection date because a result gathered before V4 Pro GA does not describe the same product snapshot as one gathered after it.
7. Measured citation behavior
Public evidence now includes one large Chinese-language observational dataset, but each study answers a different question. No current public benchmark establishes a universal DeepSeek citation rate, source preference, or ranking formula across topics, languages, dates, and surfaces.
| Evidence | Surface and method | Result that can be reused | Boundary |
|---|---|---|---|
| Tow Center, March 2025 | DeepSeek Search received one excerpt from each of 200 news articles and had to identify the original source | DeepSeek misattributed 115 of the 200 excerpts | Source-identification stress test, not ordinary user queries; predates V4 |
| Stanford FMTI, December 2025 | Disclosure audit of DeepSeek documentation and related materials | No crawler name or crawler opt-out protocol was disclosed | Transparency result, not a retrieval or citation-performance test |
| Zheng et al., August 2026 | Official consumer interface, the model displayed as DeepSeek V4 Flash, Deep Think and Web Search enabled; 52 mpox questions, one run each in Wenzhou during May 2026 | Confirms a recent, explicitly labeled search-enabled consumer setup and reports response-quality measures | One health topic, one region, no repeated runs, and no general citation-rate metric |
| Chinese-language generative search study, July 2026 | 614 Chinese queries across the web and app interfaces of four platforms, three replications; 214,119 raw records and 160,860 cleaned citation records | DeepSeek averaged 9.2 citations per answer on both surfaces; among 580 paired queries, exact-URL overlap was 40.8% and domain overlap was 50.7% | Preprint and observational June-July 2026 snapshot; no rejected candidate pool or ranking mechanism |
The Frontiers study is useful because it labels the surface, mode, location, dates, account type, and query procedure. It also warns that the displayed model name does not independently verify the hosted backend. The larger Chinese-language citation study adds repeated, cross-surface observations and makes its cleaned citation data public. Neither study reveals the rejected candidate pool or a causal ranking mechanism.
A defensible DeepSeek test should record at least five outcomes: whether search ran, which sources were displayed, whether the answer cited the correct origin, whether the target brand was mentioned, and whether the result recurred across runs. Search mode on and off should be stored separately. Chinese and English prompts should also be separate cohorts rather than pooled into one percentage.
The API can support repeated collection, but its results cannot be merged with consumer observations. Likewise, a result from Kimi, Metaso, or Baidu AI Search cannot fill a missing DeepSeek measurement. The same query and metric are necessary, but they are not sufficient unless the surface, date, language, region, and retrieval setting also match.
8. Optimizing for DeepSeek
DeepSeek-specific optimization begins with what can be tested. The company has not published a ranking formula, so every content recommendation needs a stated evidence level and an observable result.
| Action | Why it is reasonable | How to verify it |
|---|---|---|
| Publish the answer in accessible, rendered HTML | Any live retrieval path needs readable source text, even though the crawler identity is unknown | Compare server responses, logs, and consumer answers for a controlled page |
| Keep facts, dates, and original evidence together | Search answers can misattribute sources; explicit provenance makes the correct origin easier to verify | Check whether the cited URL is the original page and supports the nearby claim |
| Write self-contained passages under literal headings | A complete passage is easier to extract and attribute after retrieval | Track whether the passage or its distinctive facts appear in cited answers |
| Maintain stable Chinese and English entity descriptions | Cross-language queries may retrieve different domains or aliases | Test paired prompts and record source language and entity naming |
| Keep developer documentation versioned and current | Technical queries can change when models, modes, and API fields change | Test version-specific questions after every relevant product release |
| Repeat a fixed prompt set | One answer cannot distinguish a stable pattern from generation variance | Store several runs with the same surface, search setting, region, and date window |
Do not add a guessed DeepSeek user-agent to robots.txt. Use the AI crawler audit process to verify requests, and preserve normal access paths that your publishing policy permits. A vendor-specific allowlist becomes defensible only after the vendor publishes verifiable identifiers or your own logs establish a controlled pattern.
Multilingual GEO matters more than literal translation. Product names, company names, technical terms, dates, and version numbers should resolve to the same entity across languages. Writing for AI citation can improve passage clarity, while AI citation tracking determines whether the change affected search, citation, mention, or nothing observable.
The test protocol should compare Chinese products without importing assumptions from one to another. Tongyi and Doubao combine models and consumer assistants; Quark AI Search and Metaso present different search contexts. Their crawler disclosures, source pools, and citation interfaces require their own evidence.
9. Why DeepSeek matters for GEO
DeepSeek combines three forms of distribution: a hosted consumer search product, a programmable API with server-side web search, and open weights used in independent deployments. That breadth creates opportunities for visibility, but it also creates a measurement trap. A source cited by a self-hosted DeepSeek model with a third-party search provider has not demonstrated visibility in DeepSeek’s official consumer product.
A cross-platform comparison should use one measurement contract:
| Field | Required record |
|---|---|
| Product surface | Official consumer app, official API, self-hosted deployment, or third-party product |
| Model and mode | The displayed model name, thinking mode, and search setting |
| Query context | Exact prompt, language, region, account state, and prior conversation |
| Retrieval result | Whether search ran and which source pages were displayed or returned |
| Visibility result | Citation, link, mention, answer use without credit, or no appearance |
| Reliability | Source accuracy, support for the cited claim, variation across repeated runs, and web-app overlap |
Apply that contract to Baidu AI Search, Doubao, Kimi, Tongyi, Quark AI Search, and Metaso before drawing a China-wide conclusion. A common spreadsheet does not make the platforms identical; it makes their differences interpretable.
DeepSeek is consequential for GEO because the same model family can appear behind many products while the surrounding retrieval and interface system assigns credit to web sources. Publishers gain clearer evidence when they label the exact surface, verify source accuracy, and measure citations separately from mentions. Those habits remain useful as DeepSeek changes models, search tools, and product modes.
References
Official DeepSeek sources:
- DeepSeek V2.5: The Grand Finale (Internet Search launch, December 10, 2024)
- Introducing DeepSeek App (January 15, 2025)
- DeepSeek-V3.1-Terminus (September 22, 2025)
- DeepSeek-V4 Preview (April 24, 2026)
- DeepSeek-V4-Pro GA Release (August 13, 2026)
- Responses API reference, Responses API guide, and API change log
- DeepSeek Privacy Policy (updated February 10, 2026)
Independent evidence:
- Tow Center for Digital Journalism, AI Search Has a Citation Problem (March 6, 2025)
- Stanford Center for Research on Foundation Models, DeepSeek Transparency Report (December 2025)
- Zheng et al., Evaluating search-enabled large language model interfaces for mpox public health consultation (August 26, 2026)
- What Do Chinese-Language Generative Search Engines Cite and Surface? A Large-Scale Empirical Study and the accompanying CN-GEO Citation Dataset (July 2026)
Frequently asked questions
Does DeepSeek search the web automatically?
Does DeepSeek publish a crawler or user-agent name?
Can the DeepSeek API search the web?
Do open-weight DeepSeek models reproduce DeepSeek Search?
How should a publisher measure visibility in DeepSeek?
Related
Sources
Primary
- DeepSeek V2.5: The Grand Finale · DeepSeek · 2024-12-10
- Introducing DeepSeek App · DeepSeek · 2025-01-15
- DeepSeek-V3.1-Terminus · DeepSeek · 2025-09-22
- DeepSeek-V4 Preview: Entering the Era of Affordable Million-Token Context · DeepSeek · 2026-04-24
- DeepSeek-V4-Pro GA Release · DeepSeek · 2026-08-13
- Responses API · DeepSeek API Docs
- Responses API Guide · DeepSeek API Docs
- DeepSeek API Change Log · DeepSeek API Docs
- DeepSeek Privacy Policy · DeepSeek · 2026-02-10
- Evaluating search-enabled large language model interfaces for mpox public health consultation: a guideline-based comparative study · Frontiers in Public Health · 2026-08-26
- What Do Chinese-Language Generative Search Engines Cite and Surface? A Large-Scale Empirical Study · arXiv · 2026-07-17
- CN-GEO Citation Dataset · WENDAOstudy
Secondary
- AI Search Has a Citation Problem · Columbia Journalism Review / Tow Center for Digital Journalism
- DeepSeek Transparency Report · Stanford Center for Research on Foundation Models