Skip to content
  • Models
  • Rankings
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Favicon for CoreWeave

CoreWeave

Browse models provided by CoreWeave (Terms of Service)

17 models

Tokens processed on OpenRouter

  • Favicon for ibm-granite
    IBM: Granite 4.2 8BGranite 4.2 8B

    Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning. It supports full, low-effort, and non-thinking modes. The model supports 12 languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese.

    by ibm-graniteAug 31, 2026131K context$0.10/M input tokens$0.15/M output tokens
  • Favicon for z-ai
    Z.ai: GLM 5.3 FlashGLM 5.3 Flash

    GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

    by z-aiAug 26, 20261.05M context$0.15/M input tokens$0.50/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.8 27BQwen3.8 27B

    Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be enabled or disabled.

    by qwenAug 14, 2026262K context$0.40/M input tokens$3/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek V4 Pro 0813DeepSeek V4 Pro 0813

    DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

    by deepseekAug 12, 20261.05M context$1.31/M input tokens$3.96/M output tokens
  • Favicon for nvidia
    NVIDIA: Nemotron 3.5 LightningNemotron 3.5 Lightning

    NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for high-throughput agentic workloads and specialized tasks that benefit from domain-specific customization.

    by nvidiaAug 11, 20261M context$0.10/M input tokens$0.25/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek V4 Flash 0731DeepSeek V4 Flash 0731

    DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

    by deepseekJul 31, 20261.05M context$0.13/M input tokens$0.28/M output tokens
  • Favicon for z-ai
    Z.ai: GLM 5.2GLM 5.2

    GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

    by z-aiJun 16, 20261.05M context$0.76/M input tokens$2.42/M output tokens
  • Favicon for moonshotai
    MoonshotAI: Kimi K2.7 CodeKimi K2.7 Code

    MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts architecture that accepts text and image input, and it always operates in a thinking mode, preserving full reasoning content across multi-turn conversations. With a 256K-token context window, it targets long-horizon coding, agentic task decomposition, and multi-turn dialogue. The model activates 32B parameters out of roughly 1T total.

    by moonshotaiJun 12, 2026262K context$0.71/M input tokens$3.50/M output tokens
  • Favicon for minimax
    MiniMax: MiniMax M3MiniMax M3

    MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks. Trained as a native multimodal model on interleaved data and tuned for multi-turn, production-like collaboration via an interactive user-simulator framework, the model is oriented toward sustained, multi-step tasks rather than single-turn execution.

    by minimaxMay 31, 20261.05M context$0.23/M input tokens$0.96/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.6 35B A3BQwen3.6 35B A3B

    Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated DeltaNet linear attention with standard gated attention layers, enabling efficient inference at a fraction of the compute cost. The model supports a 262K token native context window (extensible to 1M via YaRN) and accepts text, image, and video inputs. It includes integrated thinking mode with reasoning traces preserved across multi-turn conversations, function calling, and structured output. Released under the Apache 2.0 license.

    by qwenApr 27, 2026262K context$0.25/M input tokens$1.25/M output tokens
  • Favicon for moonshotai
    MoonshotAI: Kimi K2.6Kimi K2.6

    Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and can convert prompts and visual inputs into production-ready interfaces. Its agent swarm architecture scales to hundreds of parallel sub-agents for autonomous task decomposition - delivering documents, websites, and spreadsheets in a single run without human oversight.

    by moonshotaiApr 20, 2026262K context$0.65/M input tokens$3.41/M output tokens
  • Favicon for google
    Google: Gemma 4 31BGemma 4 31B

    Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function calling, and multilingual support across 140+ languages. Strong on coding, reasoning, and document understanding tasks. Apache 2.0 license.

    by googleApr 2, 2026262K context$0.10/M input tokens$0.34/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek V3.1DeepSeek V3.1

    DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context training process, reaching up to 128K tokens, and uses FP8 microscaling for efficient inference. Users can control the reasoning behaviour with the reasoning enabled boolean. Learn more in our docs The model improves tool use, code generation, and reasoning efficiency, achieving performance comparable to DeepSeek-R1 on difficult benchmarks while responding more quickly. It supports structured tool calling, code agents, and search agents, making it suitable for research, coding, and agentic workflows. It succeeds the DeepSeek V3-0324 model and performs well on a variety of tasks.

    by deepseekAug 21, 2025131K context$0.55/M input tokens$1.65/M output tokens
  • Favicon for openai
    OpenAI: gpt-oss-120bgpt-oss-120b

    gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized to run on a single H100 GPU with native MXFP4 quantization. The model supports configurable reasoning depth, full chain-of-thought access, and native tool use, including function calling, browsing, and structured output generation.

    by openaiAug 5, 2025131K context$0.03/M input tokens$0.17/M output tokens
  • Favicon for openai
    OpenAI: gpt-oss-20bgpt-oss-20b

    gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deployability on consumer or single-GPU hardware. The model is trained in OpenAI’s Harmony response format and supports reasoning level configuration, fine-tuning, and agentic capabilities including function calling, tool use, and structured outputs.

    by openaiAug 5, 2025131K context$0.03/M input tokens$0.13/M output tokens
  • Favicon for meta-llama
    Meta: Llama 3.3 70B InstructLlama 3.3 70B Instruct

    The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model is optimized for multilingual dialogue use cases and outperforms many of the available open source and closed chat models on common industry benchmarks. Supported languages: English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai. Model Card

    by meta-llamaDec 6, 2024131K context$0.71/M input tokens$0.71/M output tokens
  • Favicon for meta-llama
    Meta: Llama 3.1 8B InstructLlama 3.1 8B Instruct

    Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to leading closed-source models in human evaluations. To read more about the model release, click here. Usage of this model is subject to Meta's Acceptable Use Policy.

    by meta-llamaJul 23, 2024131K context$0.22/M input tokens$0.22/M output tokens