Inference Race: OpenAI Cut Inference Costs in Half. AMD and Cerebras Split AI

Por Turing Post TV · 1 ago 2026 · 17:21

Visualizaciones
4.1K vistas
Likes
205 likes
Comentarios
35 comentarios

En resumen

  • OpenAI ha reducido los costos de inferencia de ChatGPT en más de la mitad.
  • AMD y Cerebras están separando las fases de inferencia de LLM para mejorar la eficiencia.
  • Se menciona un aumento de hasta 5x en tokens generados por segundo por vatio.
  • La carrera por la inferencia de IA se centra en optimizar el uso de chips existentes.
  • El clip discute cómo estas innovaciones están transformando el hardware y software de IA.

Resumen a partir del título y la descripción del video (sin análisis de comentarios).

One AI request may soon begin on one computer and finish on another. WHAT?! AMD and Cerebras are separating the two phases of LLM inference: Helios processes prompts and long context, while Cerebras generates tokens. They claim up to 5x more tokens per second per watt, although the figure is based on internal modeling. Meanwhile, OpenAI reportedly cut inference costs for one segment of ChatGPT by more than half through an undisclosed optimization. The inference race is shifting from installing more chips to extracting more useful work from them. Attention Span explains how the race to make AI inference faster and cheaper is reshaping hardware and the software that orchestrates it. 👉 Subscribe for high-signal AI mechanics 👉 Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: TuringPost.com 👉 Interviews: @realturingpost Sources and further reading The Information on OpenAI's reported inference optimization: https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half OpenAI on GPT-5.6 inference and agent-harness efficiency: https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/ AMD and Cerebras announcement: https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution AMD Helios architecture: https://www.amd.com/en/blogs/2026/amd-launches-helios-the-highest-performing-rackscale-ai-infrastructure-solution.html AMD Helios networking: https://www.amd.com/en/blogs/2026/amd-helios-resilient-scale-up-networking-for-ai.html Cerebras WSE-3: https://www.cerebras.ai/chip Cerebras CS-3 datasheet: https://cdn.sanity.io/files/e4qjo92p/production/0d73d528371618c0372fcb9de9b3c0da703adf9e.pdf Cerebras inference architecture: https://www.cerebras.ai/blog/introducing-cerebras-inference-ai-at-instant-speed Kimi K2.6 model card: https://huggingface.co/moonshotai/Kimi-K2.6 AWS and Cerebras disaggregated inference: https://www.aboutamazon.com/news/aws/aws-cerebras-ai-inference AMD and vLLM MORI-IO disaggregation test: https://vllm.ai/blog/2026-04-07-moriio-kv-connector PagedAttention: https://doi.org/10.1145/3600006.3613165 Splitwise: https://arxiv.org/abs/2311.18677 DistServe: https://arxiv.org/abs/2401.09670 TetriInfer: https://arxiv.org/abs/2401.11181 Mooncake: https://www.usenix.org/conference/fast25/presentation/qin NVIDIA Dynamo: https://www.nvidia.com/dynamo/ NVIDIA Groq 3 LPX: https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/ CDC 6600: https://www.cisl.ucar.edu/ncar-supercomputing-history/cdc6600 #ArtificialIntelligence #AI #MachineLearning #LLM #Inference #OpenAI #AMD #Cerebras #AIInfrastructure #FutureOfAI