The Archive
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Accelerating Long-Context Model Training in JAX and XLA
Large language models (LLMs) are rapidly expanding their context windows, with recent models supporting sequences of 128K tokens, 256K tokens, and beyond.... Large language models (LLMs) are rapidly expanding their context windows, with recent models supporting sequences of 128K tokens, 256K tokens, and beyond. However, training these models with extended context lengths presents significant computational and communication challenges. As context lengths grow, the memory and communication overhead of attention mechanisms scale quadratically… Source
Defining AI automation: A new kind of workplace
Enterprise AI strategy template and guidance on choosing automation approaches; generic business advisory content.
The Sora feed philosophy
OpenAI outlines Sora feed design philosophy emphasizing personalized recommendations and parental controls.
Optimizing Communication for Mixture-of-Experts Training with Hybrid Expert Parallel
In LLM training, Expert Parallel (EP) communication for hyperscale mixture-of-experts (MoE) models is challenging. EP communication is essentially all-to-all,... In LLM training, Expert Parallel (EP) communication for hyperscale mixture-of-experts (MoE) models is challenging. EP communication is essentially all-to-all, but due to its dynamics and sparseness (only topk experts per AI token instead of all experts), it’s challenging to implement and optimize. This post details an efficient MoE EP communication solution, Hybrid-EP, and its use in the… Source
Anthropic partners with Allen Institute and Howard Hughes Medical Institute to accelerate scientific discovery
Anthropic partners with Allen Institute and Howard Hughes Medical Institute to accelerate scientific discovery.
Import AI 443: Into the mist: Moltbook, agent ecologies, and the internet in transition
Import AI 443 examines agent ecology systems, Moltbook framework, and adversarial agent corruption risks.
Snowflake and OpenAI partner to bring frontier intelligence to enterprise data
OpenAI and Snowflake announce $200M partnership embedding frontier models and agents directly in Snowflake's data platform.
xAI joins SpaceX
SpaceX acquires xAI, consolidating AI development with rocket/hardware infrastructure.
Introducing the Codex app
OpenAI releases Codex app for macOS enabling parallel multi-agent coding workflows with long-running task support.
Advancing GPU Programming with the CUDA Tile IR Backend for OpenAI Triton
NVIDIA CUDA Tile is a GPU-based programming model that targets portability for NVIDIA Tensor Cores, unlocking peak GPU performance. One of the great things... NVIDIA CUDA Tile is a GPU-based programming model that targets portability for NVIDIA Tensor Cores, unlocking peak GPU performance. One of the great things about CUDA Tile is that you can build your own DSL on top of it. This post shares the work NVIDIA is doing to integrate CUDA Tile as a backend for OpenAI Triton, an open source Python DSL designed to write DL kernels for GPUs. Source
Establishing a Scalable Sparse Ecosystem with the Universal Sparse Tensor
Sparse tensors are vectors, matrices, and higher-dimensional generalizations with many zeros. They are crucial in various fields such as scientific computing,... Sparse tensors are vectors, matrices, and higher-dimensional generalizations with many zeros. They are crucial in various fields such as scientific computing, signal processing, and deep learning due to their efficiency in storage, computation, and power. Despite their benefits, handling sparse tensors manually or through existing libraries is often cumbersome, error-prone, nonportable… Source
Practical Security Guidance for Sandboxing Agentic Workflows and Managing Execution Risk
AI coding agents enable developers to work faster by streamlining tasks and driving automated, test-driven development. However, they also introduce a... AI coding agents enable developers to work faster by streamlining tasks and driving automated, test-driven development. However, they also introduce a significant, often overlooked, attack surface by running tools from the command line with the same permissions and entitlements as the user, making them computer use agents, with all the risks those entail. The primary threat to these tools is… Source
Project Genie: Experimenting with infinite, interactive worlds
Project Genie lets Google AI Ultra subscribers create and explore infinite interactive worlds via experimental prototype.
Inside OpenAI’s in-house data agent
OpenAI describes internal data agent using GPT-5 and Codex with memory for reasoning over large datasets.
Retiring GPT-4o, GPT-4.1, GPT-4.1 mini, and OpenAI o4-mini in ChatGPT
OpenAI announces retirement of GPT-4o, GPT-4.1, and o4-mini from ChatGPT effective February 13, 2026; API unaffected.
Taisei Corporation shapes the next generation of talent with AI
Taisei Corporation deploys ChatGPT Enterprise for internal HR and talent development workflows.
ServiceNow chooses Claude to power customer apps and increase internal productivity
ServiceNow adopts Claude to power customer-facing apps and boost internal productivity.
Ensuring Balanced GPU Allocation in Kubernetes Clusters with Time-Based Fairshare
NVIDIA Run:ai v2.24 introduces time-based fairshare, a new scheduling mode that brings fair-share scheduling with time awareness for over-quota resources to... NVIDIA Run:ai v2.24 introduces time-based fairshare, a new scheduling mode that brings fair-share scheduling with time awareness for over-quota resources to Kubernetes clusters. This capability, built on the open source KAI Scheduler that powers NVIDIA Run:ai, addresses a long-standing challenge in shared GPU infrastructure. Consider two teams with equal priority sharing a cluster. Source
Speeding Up Variable-Length Training with Dynamic Context Parallelism and NVIDIA Megatron Core
This post introduces Dynamic Context Parallelism (Dynamic-CP), a scheduling approach in NVIDIA Megatron Core used for LLM post-training or DiT pre-training. It... This post introduces Dynamic Context Parallelism (Dynamic-CP), a scheduling approach in NVIDIA Megatron Core used for LLM post-training or DiT pre-training. It dynamically selects the CP size per microbatch to efficiently handle variable-length sequences, achieving up to 1.48x speedup on real-world datasets. In large-scale model training, an often-overlooked bottleneck arises from the… Source
Updating Classifier Evasion for Vision Language Models
Advances in AI architectures have unlocked multimodal functionality, enabling transformer models to process multiple forms of data in the same context. For... Advances in AI architectures have unlocked multimodal functionality, enabling transformer models to process multiple forms of data in the same context. For instance, vision language models (VLMs) can generate output from combined image and text input, enabling developers to build systems that interpret graphs, process camera feeds, or operate with traditionally human interfaces like desktop… Source
EMEA Youth & Wellbeing Grant
OpenAI launches €500K EMEA youth safety and wellbeing grant program for NGOs and researchers.
The next chapter for AI in the EU
OpenAI publishes EU Economic Blueprint 2.0 highlighting AI adoption initiatives and partnerships across Europe.
Grok Imagine API
Grok Imagine API offers video generation with stated advances in quality, cost, latency.
Keeping your data safe when an AI agent clicks a link
OpenAI details safeguards protecting agent data when clicking links, preventing URL exfiltration and prompt injection.
Accelerating Diffusion Models with an Open, Plug-and-Play Offering
Recent advances in large-scale diffusion models have revolutionized generative AI across multiple domains, from image synthesis to audio generation, 3D asset... Recent advances in large-scale diffusion models have revolutionized generative AI across multiple domains, from image synthesis to audio generation, 3D asset creation, molecular design, and beyond. These models have demonstrated unprecedented capabilities in producing high-quality, diverse outputs across various conditional generation tasks. Despite these successes… Source