Skip to content

Ling 3.0 Flash

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*.

Per 1M tokens
$0.0221 · $0.0662
Context
262Ktokens
Supports
  • Calls tools
  • Thinks before it answers
opencode.json
{
  "provider": {
    "deference": {
      "npm": "@ai-sdk/openai-compatible",
      "options": { "baseURL": "https://deference.si/v1", "apiKey": "{env:DEFERENCE_API_KEY}" },
      "models": { "inclusionai/ling-3.0-flash": { "limit": { "context": 262144, "output": 32768 } } }
    }
  }
}
Needs an API keyGet one

Pricing and limits

  • Cache read$0.0044per 1M tokens
  • Max output32,768tokens

OpenRouter's price plus OpenRouter's 5% platform fee. Deference adds no markup.