Dashboard

MoonshotAI: Kimi Linear 48B A3B Instruct - NeuralHub | NeuralHub

MoonshotAI: Kimi Linear 48B A3B Instruct

Model Details

Company: moonshotai

Created: 12/11/2025

Description

Kimi Linear is a hybrid linear attention architecture that outperforms traditional full attention methods across various contexts, including short, long, and reinforcement learning (RL) scaling regimes. At its core is Kimi Delta Attention (KDA)—a refined version of Gated DeltaNet that introduces a more efficient gating mechanism to optimize the use of finite-state RNN memory. Kimi Linear achieves superior performance and hardware efficiency, especially for long-context tasks. It reduces the need for large KV caches by up to 75% and boosts decoding throughput by up to 6x for contexts as long as 1M tokens.

Technical Specifications

Context Window

1049k tokens

Max Output

1049k tokens

Pricing (Input / Output)

$0.0007 / $0.0009 per 1M

Architecture

transformer

Modality

text->text

API Usage

Example API Call

curl -X POST https://api.neuralhub.xyz/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer NEURALHUB_API_KEY" \
-d '{
  "model": "moonshotai/kimi-linear-48b-a3b-instruct",
  "messages": [
    { "role": "system", "content": "You are a helpful assistant." },
    { "role": "user", "content": "" }
  ],
  "temperature": 0.7,
  "max_tokens": 500,
  "top_p": 0.9
}'

Response Format

The API returns an OpenAI-compatible response. Example:

{
  "id": "chatcmpl-<uuid>",
  "object": "chat.completion",
  "created": 1765590421,
  "model": "moonshotai/kimi-linear-48b-a3b-instruct",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "The answer to life, the universe, and everything is famously 42..."
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 26,
    "completion_tokens": 169,
    "total_tokens": 195
  }
}