← Материалы разборы Yersham
разбор · статьятема-фронтир · H 1.58
7.0
из 10
смотреть по главам — взять две-три нужные минуты
оценка машинная и частично зависит от длины ролика — спорите, открывайте оригинал

DeepSeek V4: кэш и off-peak снижают цену в разы

Models Pricing | DeepSeek API Docs

DeepSeek V4-Flash и V4-Pro с контекстом 1M, дешевле в off-peak, кэш-скидки, поддержка OpenAI и Anthropic API.

DeepSeek V4-Flash и V4-Pro с контекстом 1M, дешевле в off-peak, кэш-скидки, поддержка OpenAI и Anthropic API.

что из этого моё

Для оркестрации моделей важно: можно маршрутизировать запросы на DeepSeek в off-peak для экономии, использовать кэш для повторяющихся промптов, а также задействовать Responses API и Anthropic API для совместимости. Контекст 1M позволяет обрабатывать большие документы, что полезно для агентов.

Что забрать
отметь, что берёшь в работу → или отбрось как не своёмоё →
переключить повторяющиеся запросы на DeepSeek Flash с кэшем, чтобы снизить стоимость в 30 раз
настроить маршрутизацию запросов на DeepSeek в off-peak часы для экономии
переключить FIM Completion в non-thinking режим для автодополнения кода
расшифровка ролика ↓

Models Pricing | DeepSeek API Docs

Skip to main content On this page Models Pricing The prices listed below are in units of per 1M tokens. A token, the smallest unit of text that the model recognizes, can be a word, a number, or even a punctuation mark. We will bill based on the total number of input and output tokens by the model.

Model Details ​ MODEL deepseek-v4-flash deepseek-v4-pro BASE URL (OpenAI Format) https://api.deepseek.com BASE URL (Anthropic Format) https://api.deepseek.com/anthropic MODEL VERSION DeepSeek-V4-Flash-0731 DeepSeek-V4-Pro-0813 THINKING MODE Supports both non-thinking and thinking (default) modes See Thinking Mode for how to switch CONTEXT LENGTH 1M MAX OUTPUT MAXIMUM: 384K FEATURES Json Output ✓ ✓ Tool Calls ✓ ✓ Responses API ✓ ✓ Anthropic API ✓ ✓ Chat Prefix Completion(Beta) ✓ ✓ FIM Completion(Beta) Non-thinking mode only Non-thinking mode only PRICING (1) 1M INPUT TOKENS (CACHE HIT) OFF-PEAK $0.007 $0.022 PEAK $0.014 $0.044 1M INPUT TOKENS (CACHE MISS) OFF-PEAK $0.22 $0.66 PEAK $0.44 $1.32 1M OUTPUT TOKENS OFF-PEAK $0.66 $1.98 PEAK $1.32 $3.96 Concurrency Limit (2) 2500 500 (1) Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC (all other hours are off-peak). (2) For more details on concurrency limits, please refer to Rate Limit Isolation

Deduction Rules ​ The expense = number of tokens × price. The corresponding fees will be directly deducted from your topped-up balance or granted balance, with a preference for using the granted balance first when both balances are available. Product prices may vary and DeepSeek reserves the right to adjust them. We recommend topping up based on your actual usage and regularly checking this page for the most recent pricing information. Model Details Deduction Rules

дальше в дело
Собрать это в маршрут
все маршруты →
не хочешь разбираться сам
Сделаю это под задачу
форматы и цены →
ещё разборы