Pricing
Start free. Scale with your token volume. Contact us for enterprise arrangements.
Explorer
For evaluation and early integration work
1 million tokens per month
- API access to parallel decoding endpoint
- Standard rate limits
- Documentation and SDK
- Community support
Build
For production workloads in active development
20 million tokens per month
- Everything in Explorer
- Higher rate limits
- Overage tokens available at usage rate
- Priority email support
- Usage dashboard and reporting
Scale
For teams with high-volume or mission-critical requirements
Volume commitment, negotiated rate
- Everything in Build
- Dedicated capacity allocation
- SLA with uptime commitment
- Custom integration support
- Dedicated account contact
Frequently asked questions
-
We count output tokens only: the tokens Inception generates per request. Input (prompt) tokens are not billed separately. Token counting follows the same tokenizer conventions as the underlying model.
-
Diffusion-style decoding refines all positions jointly rather than committing left to right. For tasks like document generation and code, the global coherence can improve because late positions can inform early ones during refinement. Outputs are not identical to autoregressive and we recommend evaluating both approaches on your specific task.
-
Explorer accounts are rate-limited at the monthly cap. Build plan customers can continue beyond the 20M included tokens at a per-token overage rate. Scale plan customers negotiate a committed volume with no hard cap during the commitment period.
-
Scale pricing is negotiated individually based on volume, use case, and infrastructure requirements. There is typically a minimum monthly commitment, but the specifics depend on your workload. Reach out via the contact form to discuss.
-
Monthly token allocations reset at the start of each billing cycle and do not carry over. This applies to both the Explorer free tier and the Build plan's included 20M tokens.