Cohere's North Mini Code Megakernel Serving Engine
Cohere releases North Mini, a decode megakernel LLM serving engine achieving 1.58× speedup over vLLM in production.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Cohere releases North Mini, a decode megakernel LLM serving engine achieving 1.58× speedup over vLLM in production.
Activation steering in LLMs validates whether steering vectors encode coherent human-value geometry using Schwartz's moral theory.
Knowledge Space Theory framework evaluates whether LLMs exhibit structured, prerequisite-dependent knowledge in mathematical reasoning.
Cohere releases Agentic Task Ecosystem dataset: 696k AI tools analyzed, only 2.6% deliver true work automation.
Dutch Books framework reveals probabilistic coherence failures in LLM forecasts via linear programming arbitrage detection.
Cohere publishes guide on small language model deployment, enterprise adoption patterns, and SLM strategy considerations.
Cohere overview of enterprise generative AI adoption patterns, benefits, and implementation challenges.
Cohere Parse converts unstructured documents and images into structured data for enterprise AI systems.
Code World Model combines LLMs and video models to separate world dynamics from visual rendering for coherent, open-ended environment simulation.
The startup came out of stealth with $6 million in seed funding and a plan to use LLMs and cybersecurity know-how to make AI queries coherent.
Statistical inference method for dictionary learning under calibration uncertainty; niche signal-processing contribution.
Cohere argues cultural awareness must be embedded in AI design beyond multilingual support to serve global users effectively.
KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history. Under aggressive budgets, this not only lowers accuracy but can also cause runaway degeneration, where the model produces incoherent or repetitive tokens until reaching the length limit. We characterize much of this loss as an information gapf caused by missing context, rather than a capability gap caused by limited model capacity. An evicted 7B model and a full-context 1.5B model make complementary errors, and an oracle choice between their ans...
Health misinformation circulating during pandemics can gain traction rapidly, creating harmful narratives that compete with public health guidance. Most topic-modelling pipelines treat engagement as an external outcome, limiting their ability to prioritise semantically coherent topics that are also rapidly diffusing. We introduce BERTopic-VP, a virality-prioritised topic-modelling framework that combines contextual embedding-based clustering (BERTopic) with a post hoc Virality Prioritisation (VP) layer. The pipeline is complemented by a two-stage hybrid misinformation detection module that fu...
Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branc...
Scientific data often describe entities whose features are jointly governed by the laws of physics, yet existing self-supervised learning (SSL) objectives largely ignore this physical coherence. We introduce imposter, a discriminative pretext task that replaces subsets of an entity's features with real observations donated by another entity and trains the encoder to identify the swapped features. Because every donated value is individually plausible, the task can only be solved by learning cross-feature physical dependencies. We evaluate the proposed objectives on global ERA5-Land reanalysis ...
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and a machine-checked proof of its specification, offers a stronger path toward trustworthy AI-generated software. Existing benchmarks in this direction either focus on individual functions or only evaluate proof generation with provided implementations. It is still an open question whether agents can make coherent implementation and proof choices across real multi-module codebases. To bridge this ga...
Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts a...
Cross-Domain Sequential Recommendation (CDSR) aims to alleviate data sparsity by transferring dynamic user interests across related domains. A key challenge lies in effectively bridging these domains. In single-domain modeling, models cannot distinguish between domain-specific and domain-invariant interests. Recent methods merge domain-specific sequences chronologically into a mixed-domain sequence to capture domain-invariant knowledge. However, they typically deploy separate encoders for the mixed-domain sequence and train them with per-domain loss aggregation. This workflow magnifies inter-...
Cohere and the University of Waterloo launch partnership to strengthen Canada’s AI talent pipeline.
Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour. Large language models (LLMs) can reason about such concepts, yet are not calibrated probability engines. We combine both components through an auditable information harness. This paper documents four iterations: V1, a dynamic score-driven Dixon-Coles baseline; V2, which maps LLM contextual ...
Cohere signs EU AI Content Transparency Code, demonstrating early compliance with EU AI Act requirements.
Expert Subspace Separation Index disentangles route coherence from candidate quality in sparse MoE routing using controlled interventions.
Cohere overview of AI applications in manufacturing: use cases and implementation considerations for industrial operators.
Hierarchical spatio-temporal transformer for multi-level emergency department demand forecasting; healthcare application, not core AI frontier.
MDTransformer uses mode-division photonic hardware with inverse-designed crossbars to accelerate Transformer inference with reduced phase-shifter overhead.
Cohere publishes explainer on agentic AI definitions, applications, and enterprise deployment considerations.
Case study: Cohere's North AI agents reduce friction for wealth managers through task automation and time savings.
Cohere launches North Automations for orchestrating multi-agent enterprise workflows with monitoring and control.
Cohere partners with Government of Canada to develop sovereign AI capabilities for public-sector services.