Today, we are excited to share our collaboration with Upstage, one of South Korea’s leading AI companies, and explore how Cerebras Inference can bring ultra-fast, production-ready AI experiences to more users around the world. Running on the Cerebras wafer-scale engine, Upstage’s Solar 31B can achieve inference speeds of up to 2,000 tokens per second.
Upstage has built a strong portfolio of enterprise AI products spanning large-scale language models, document intelligence, and AI workflow automation. The company’s Solar model family is designed to deliver enterprise-grade linguistic intelligence with a focus on speed, rationale, and practical deployment across business workflows.
“We are excited to work with Cerebras to provide developers with solar models with industry-leading inference speeds. Together, we will make it easy to build fast, production-ready AI applications at scale,” said Sung Kim, CEO of Upstage.
By running Upstage’s Solar model on Cerebras Inference Cloud, developers can achieve industry-leading inference speeds without changing their existing application development workflows.
Julie Choi, Chief Marketing Officer of Cerebras, said, “Korea is at the forefront of AI innovation, and we are excited to partner with Upstage to help developers and businesses in Korea and abroad achieve new levels of inference performance.”
Expanding high-speed inference to Korea
Upstage builds AI for mission-critical enterprise workflows across major industries in Korea, including finance, insurance, healthcare, and manufacturing. Cerebras provides an inference infrastructure that allows advanced AI applications to run faster, respond in real-time, and easily scale.
Together, we bring lightning-fast inference to one of the world’s most advanced AI markets, enabling developers to build multilingual real-time AI applications with industry-leading performance.
Build a real-time AI experience
Modern AI applications increasingly capture information, reason across contexts, produce structured output, and interact with users in real time.
Cerebras Inference is designed for exactly this type of workload. With OpenAI-compatible APIs, developers can start building on Cerebras with minimal code changes while taking advantage of the speed of wafer-scale computing. Cerebras provides production-scale inference for enterprises with dedicated support, custom model capabilities, and an infrastructure designed for high-volume deployments.
Below is an example of Solar 31b retrieving 246 sources and running a detailed research query. An equivalent query with a similar number of sources in Sonnet 4.6 takes 8 minutes.
By combining Upstage’s model and application expertise with Cerebras’ ultra-fast inference infrastructure, we are working towards a future where powerful AI applications are not only more capable but also dramatically more responsive.
