we heart startups

+1 202 555 0180

Have a question, comment, or concern? Our dedicated team of experts is ready to hear and assist you. Reach us through our social media, phone, or live chat.

Meta and Cerebras Systems introduce Llama API achieving 2,600 tokens per second, outperforming GPU solutions by up to 18x, as Meta seeks to commercialize Llama models and compete with OpenAI and Google in AI inference.

Meta partners with Cerebras Systems to launch a Llama API delivering 2,600 tokens/second, exceeding GPU solutions by up to 18x. Transforming Llama models into a commercial service, Meta aims to challenge OpenAI and Google in AI inference.

Source: venturebeat.com

Prev Post

Freepik debuts F Lite, an AI image generator with 10B parameters, trained on 80M licensed images using 64 Nvidia H100 GPUs; developed with Fal.ai, it offers standard and texture versions, similar to Adobe and Shutterstock’s AI tools.

Next Post

Meta’s Llama AI models hit 1.2B downloads, up from 1B in March, now serving nearly 1B users; Chris Cox highlights strong developer interest at LlamaCon amid competition from Alibaba’s Qwen3 models

Leave a Reply
Read next