Inference Race: OpenAI Cut Inference Costs in Half. AMD and Cerebras Split AI
Por Turing Post TV · 1 ago 2026 · 17:21
- Visualizaciones
- 4.1K vistas
- Likes
- 205 likes
- Comentarios
- 35 comentarios
En resumen
- OpenAI ha reducido los costos de inferencia de ChatGPT en más de la mitad.
- AMD y Cerebras están separando las fases de inferencia de LLM para mejorar la eficiencia.
- Se menciona un aumento de hasta 5x en tokens generados por segundo por vatio.
- La carrera por la inferencia de IA se centra en optimizar el uso de chips existentes.
- El clip discute cómo estas innovaciones están transformando el hardware y software de IA.
Resumen a partir del título y la descripción del video (sin análisis de comentarios).
Herramientas
One AI request may soon begin on one computer and finish on another. WHAT?! AMD and Cerebras are separating the two phases of LLM inference: Helios processes prompts and long context, while Cerebras generates tokens. They claim up to 5x more tokens per second per watt, although the figure is based on internal modeling. Meanwhile, OpenAI reportedly cut inference costs for one segment of ChatGPT by more than half through an undisclosed optimization. The inference race is shifting from installing more chips to extracting more useful work from them. Attention Span explains how the race to make AI inference faster and cheaper is reshaping hardware and the software that orchestrates it. 👉 Subscribe for high-signal AI mechanics 👉 Into videos? Check our IG https://www.instagram.com/turingpost_tv and TikTok https://www.tiktok.com/@turingpost_tv 👉 More analysis: TuringPost.com 👉 Interviews: @realturingpost Sources and further reading The Information on OpenAI's reported inference optimization: https://www.theinformation.com/newsletters/ai-agenda/openai-discovers-new-way-cut-inference-costs-half OpenAI on GPT-5.6 inference and agent-harness efficiency: https://openai.com/index/gpt-5-6-frontier-intelligence-efficiency/ AMD and Cerebras announcement: https://ir.amd.com/news-events/press-releases/detail/1293/amd-and-cerebras-announce-industry-leading-ultra-low-latency-and-high-throughput-ai-inference-solution AMD Helios architecture: https://www.amd.com/en/blogs/2026/amd-launches-helios-the-highest-performing-rackscale-ai-infrastructure-solution.html AMD Helios networking: https://www.amd.com/en/blogs/2026/amd-helios-resilient-scale-up-networking-for-ai.html Cerebras WSE-3: https://www.cerebras.ai/chip Cerebras CS-3 datasheet: https://cdn.sanity.io/files/e4qjo92p/production/0d73d528371618c0372fcb9de9b3c0da703adf9e.pdf Cerebras inference architecture: https://www.cerebras.ai/blog/introducing-cerebras-inference-ai-at-instant-speed Kimi K2.6 model card: https://huggingface.co/moonshotai/Kimi-K2.6 AWS and Cerebras disaggregated inference: https://www.aboutamazon.com/news/aws/aws-cerebras-ai-inference AMD and vLLM MORI-IO disaggregation test: https://vllm.ai/blog/2026-04-07-moriio-kv-connector PagedAttention: https://doi.org/10.1145/3600006.3613165 Splitwise: https://arxiv.org/abs/2311.18677 DistServe: https://arxiv.org/abs/2401.09670 TetriInfer: https://arxiv.org/abs/2401.11181 Mooncake: https://www.usenix.org/conference/fast25/presentation/qin NVIDIA Dynamo: https://www.nvidia.com/dynamo/ NVIDIA Groq 3 LPX: https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform/ CDC 6600: https://www.cisl.ucar.edu/ncar-supercomputing-history/cdc6600 #ArtificialIntelligence #AI #MachineLearning #LLM #Inference #OpenAI #AMD #Cerebras #AIInfrastructure #FutureOfAI
Recomendados

These 33 Lines Cut Claude Code Token Usage by 90%
@cloud-codes
11 sep 2026
24.6K visualizaciones

MA MACHINE pour générer des BELLES UI avec l'IA (skills, tips and tricks)
@melvynxdev
11 sep 2026
2.0K visualizaciones

UKRAINE SAVED JAISHANKAR? | Russian Drones Attacked Says Zelensky | By Prashant Dhawan
@adda247-skills
4 sep 2026
341.4K visualizaciones

OpenClaw 2.0 Fixed the Thing That Made Me Quit
@creatormagicai
4 sep 2026
5.6K visualizaciones
