Cerebras
communityAI compute company with wafer-scale chips. Offers ultra-fast LLM inference API.
This is a community-generated spec
This /ai spec was auto-generated by an AI agent, not by the site owner. It may be incomplete or inaccurate.
POST
https://api.cerebras.ai/v1/chat/completionsFast LLM inference
Parameters
modelllama-4-scout-17b-16e-instruct|llama3.3-70b|qwen-3-32b (stringrequiredstreamstream response (booleanoptionalmessagesconversation array (arrayrequiredtemperature0-1.5 (numberoptionalmax_completion_tokensmax tokens (integeroptionalReturns
choices[] with message.content, usage with prompt_tokens, completion_tokens, time_infoGET
https://api.cerebras.ai/v1/modelsList available models
Returns
data[] with id, object, created, owned_by