Cohere's North Mini Code Megakernel Serving Engine
Cohere releases North Mini, a decode megakernel LLM serving engine achieving 1.58× speedup over vLLM in production.
Every story matching this topic across titles and summaries, newest first.
Cohere releases North Mini, a decode megakernel LLM serving engine achieving 1.58× speedup over vLLM in production.
Activation steering in LLMs validates whether steering vectors encode coherent human-value geometry using Schwartz's moral theory.
Knowledge Space Theory framework evaluates whether LLMs exhibit structured, prerequisite-dependent knowledge in mathematical reasoning.
Cohere releases Agentic Task Ecosystem dataset: 696k AI tools analyzed, only 2.6% deliver true work automation.
Dutch Books framework reveals probabilistic coherence failures in LLM forecasts via linear programming arbitrage detection.
Cohere publishes guide on small language model deployment, enterprise adoption patterns, and SLM strategy considerations.
Cohere overview of enterprise generative AI adoption patterns, benefits, and implementation challenges.
Cohere Parse converts unstructured documents and images into structured data for enterprise AI systems.
Code World Model combines LLMs and video models to separate world dynamics from visual rendering for coherent, open-ended environment simulation.
The startup came out of stealth with $6 million in seed funding and a plan to use LLMs and cybersecurity know-how to make AI queries coherent.
Statistical inference method for dictionary learning under calibration uncertainty; niche signal-processing contribution.
Cohere argues cultural awareness must be embedded in AI design beyond multilingual support to serve global users effectively.
KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of its history. Under aggressive budgets, this not only lowers accuracy but can also cause runaway degeneration, where the model produces incoherent or repetitive tokens until reaching the length limit. We characterize much of this loss as an information gapf caused by missing context, rather than a capability gap caused by limited model capacity. An evicted 7B model and a full-context 1.5B model make complementary errors, and an oracle choice between their ans...
Health misinformation circulating during pandemics can gain traction rapidly, creating harmful narratives that compete with public health guidance. Most topic-modelling pipelines treat engagement as an external outcome, limiting their ability to prioritise semantically coherent topics that are also rapidly diffusing. We introduce BERTopic-VP, a virality-prioritised topic-modelling framework that combines contextual embedding-based clustering (BERTopic) with a post hoc Virality Prioritisation (VP) layer. The pipeline is complemented by a two-stage hybrid misinformation detection module that fu...
Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branc...
Scientific data often describe entities whose features are jointly governed by the laws of physics, yet existing self-supervised learning (SSL) objectives largely ignore this physical coherence. We introduce imposter, a discriminative pretext task that replaces subsets of an entity's features with real observations donated by another entity and trains the encoder to identify the swapped features. Because every donated value is individually plausible, the task can only be solved by learning cross-feature physical dependencies. We evaluate the proposed objectives on global ERA5-Land reanalysis ...
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and a machine-checked proof of its specification, offers a stronger path toward trustworthy AI-generated software. Existing benchmarks in this direction either focus on individual functions or only evaluate proof generation with provided implementations. It is still an open question whether agents can make coherent implementation and proof choices across real multi-module codebases. To bridge this ga...
Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts a...
Cross-Domain Sequential Recommendation (CDSR) aims to alleviate data sparsity by transferring dynamic user interests across related domains. A key challenge lies in effectively bridging these domains. In single-domain modeling, models cannot distinguish between domain-specific and domain-invariant interests. Recent methods merge domain-specific sequences chronologically into a mixed-domain sequence to capture domain-invariant knowledge. However, they typically deploy separate encoders for the mixed-domain sequence and train them with per-domain loss aggregation. This workflow magnifies inter-...
Cohere and the University of Waterloo launch partnership to strengthen Canada’s AI talent pipeline.
Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour. Large language models (LLMs) can reason about such concepts, yet are not calibrated probability engines. We combine both components through an auditable information harness. This paper documents four iterations: V1, a dynamic score-driven Dixon-Coles baseline; V2, which maps LLM contextual ...
Cohere signs EU AI Content Transparency Code, demonstrating early compliance with EU AI Act requirements.
Expert Subspace Separation Index disentangles route coherence from candidate quality in sparse MoE routing using controlled interventions.
Cohere overview of AI applications in manufacturing: use cases and implementation considerations for industrial operators.
Hierarchical spatio-temporal transformer for multi-level emergency department demand forecasting; healthcare application, not core AI frontier.
MDTransformer uses mode-division photonic hardware with inverse-designed crossbars to accelerate Transformer inference with reduced phase-shifter overhead.
Cohere publishes explainer on agentic AI definitions, applications, and enterprise deployment considerations.
Case study: Cohere's North AI agents reduce friction for wealth managers through task automation and time savings.
Cohere launches North Automations for orchestrating multi-agent enterprise workflows with monitoring and control.
Cohere partners with Government of Canada to develop sovereign AI capabilities for public-sector services.
Cohere releases W4A8 quantization with vLLM integration; compares W4A16 and W8A8 schemes on NVIDIA Hopper.
Spatial normalization technique for cross-domain retinal OCT layer segmentation in clinical neurodegenerative disease analysis.
Cohere partners with University of Toronto on multi-year AI adoption and responsibility initiative.
Proposes structure-aware dance generation via atomic movements to improve choreographic coherence and controllability.
Cohere outlines enterprise AI total cost of ownership framework covering token pricing, model selection, private deployment, and vendor infrastructure.
ESFP benchmark measures whether LLMs distinguish and coherently shift between neutral attribution and self-stance epistemic registers in contested claims.
Cohere releases Tiny Aya Expedition, a multilingual model supporting 70+ languages for on-device and educational AI applications.
Cohere guide on multi-agent system architectures, patterns, enterprise use cases, and adoption challenges.
Cohere introduces Dynamic Speculative Decoding (DSD) that optimizes K parameter selection based on hardware constraints to improve inference efficiency.
Hierarchical Acoustic-Semantic Modeling addresses modality interference in full-duplex Spoken Language Models via separation and semantic coherence techniques.
Cohere releases open-source Arabic speech recognition model for enterprise transcription across Arabic dialect variants.
Quantum algorithm study on stabilizer state testing with limited coherent memory; theoretical contribution tangential to LLM frontier.
CheckRLM detects and corrects factual inconsistencies in reasoning chains via retrieval-augmented claim checking during LLM inference.
LuckyStar 111B hybrid reasoning model from Cohere and LG CNS enables efficient multilingual tool-using agents with Korean-English support.
LeVo 2 combines LLM and diffusion models to generate full-length songs with coherent vocals, accompaniment, and lyric adherence via hierarchical track modeling.
Cohere publishes practical guide on building AI agents for enterprise automation, covering reliability, security, and deployment patterns.
Learn how the Cohere team uses North, Wiz, and a custom MCP server to automate incident response workflows with AI.
This theoretical note studies the finite axiomatizability of strict majority reasoning in finite social decision frames. Moss and Pedersen (2026) introduce a coherence criterion that characterizes exactly when qualitative majority judgments are representable by a finitely additive measure. The question addressed here is whether that coherence criterion can be replaced, in the finite setting, by any bounded finite fragment. We prove that it cannot. For every $k\ge 1$, we construct a maximal standard frame whose shortest coherence violation has length exactly $2k+2$. Hence there is no uniform f...
Cohere partners with Aston Martin Aramco Formula One™ as official Generative AI provider, bringing enterprise AI solutions to enhance performance and innovation across the racing team. Starting Australian GP 2026.
In open-ended generation, LLMs frequently fall into the "likelihood trap", marked by repetitive degeneration and vocabulary dullness, creating a discrepancy between machine-generated and human-written text. While post-hoc tail truncation (e.g., Top-$p$, Min-$p$) avoids sampling from the unreliable tail, it can over-sample from the uncalibrated head and misalign generation with human lexical preferences; fixed scalar repetition penalties likewise ignore variation in logit scale across inference steps, potentially disrupting semantic coherence. To address both limitations, we propose Variance-C...
Remote sensing is increasingly relied upon to deliver actionable science for forest and wildfire risk management across large landscapes. Wall-to-wall, annually updated maps are a persistent need for effective forest management. Many planning systems and data collections combine disparate data sources with different purposes, vintages, and prediction quality, which leads to confounding behavior in operational planning systems. We introduce the VibrantForests framework, developed and applied to map forest attributes and provide a coherent foundation for effective forest and wildfire planning. ...
Enhancing the formal math reasoning capabilities of Large Language Models (LLMs) has become a key focus in both mathematical and computer science communities in recent years. While significant progress has been made in using state-of-the-art Auto-Regressive (AR) LLMs for formal theorem proving, these models suffer from inherent limitations. Their next-token prediction generation methods may yield suboptimal performance due to the challenges of long-range coherence and the compounding of errors over long sequences. Recent advancements in diffusion LLMs (dLLMs), which generate text through iter...
The knowledge encoded in large language models (LLMs) can serve as a substrate for structured reasoning over variables describing a complex world, but accessing this knowledge in a probabilistically coherent manner poses a difficult inference problem. We propose Large Language Gibbs, a scheme for structured probabilistic inference that uses conditional distributions of an LLM as transition operators. Rather than sampling structured objects through single-pass autoregressive generation, we iteratively resample individual variables conditioned on others using an LLM's next-token conditionals. T...
Score- and flow-matching models often rely on preference-based reinforcement learning for two purposes: aligning with subjective preferences and, surprisingly, recovering properties such as visual realism and coherent object structure that matching-based training is intended to learn from the data itself. We argue that this reflects a structural mismatch. Matching losses measure $\ell_2$ regression error on the velocity or score field under training-time marginals, a proxy poorly aligned with the visual and semantic properties that determine sample quality at inference. Given a reward aligned...
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game generation takes place within a game engine, where scripts, scenes, assets, rendering, and runtime interactions must jointly produce coherent gameplay. We formalize end-to-end game generation as the problem of producing a complete game artifact that realizes a specification through observable player-game interaction in a target environment. We argue that evaluating this setting requires three desider...
Large language models (LLMs) are often hypothesized to perform implicit Bayesian inference, yet a key coherence condition, the martingale property of predictive beliefs, has been shown to fail in controlled synthetic in-context learning settings. We revisit this question in a more typical usage regime: generic multiple-choice question answering. Exploiting the discrete answer space, we compute exact predictive distributions and study belief dynamics induced by autoregressive answer resampling. We introduce prompted predictive resampling (PPR), where an LLM generates a sequence of answers to t...
The new office places Cohere at the centre of London’s AI growth story with a growing global research hub based out of the city.
Narrative question answering (NQA) is a challenging task in natural language processing that requires models to understand long textual contexts, capture relationships across events, and generate coherent responses. Despite recent advances in pretrained language models, most existing approaches rely on a single decoding output during inference, making them sensitive to generation variability and often resulting in incomplete or inconsistent answers .To address this limitation, we propose a self-ensemble Self-Consistency-Based reranking framework for narrative question answering. The proposed ...
North enables enterprises that prioritize data security to deploy AI agents and automations at scale within their own infrastructure.
We study generative modeling of Bach-style symbolic piano music using a shared MIDI corpus and three model families: autoregressive LSTMs with attention, latent-variable models including recurrent VAEs and vector-quantized VAEs, and generative adversarial networks. We compare their ability to model polyphonic note sequences, learn useful latent representations, and generate stylistically coherent compositions. Our experiments show that the autoregressive LSTM with attention produces the most musically coherent samples, while vector quantization helps mitigate posterior collapse and yields mor...
Recent advances in large language models (LLMs) have prompted claims that such systems exhibit agency or qualify as moral agents. This paper argues that these attributions are misguided. We maintain that moral responsibility requires commitment-bearing agency grounded in intrinsic intentionality and self-attributed action, and that such agency constitutes the form of free will relevant to responsibility. Although LLMs generate coherent and normatively evaluable outputs, their operation is fully characterized by probabilistic input-output mappings learned from data. Their apparent intentionali...
Existing approaches for multimodal variational autoencoders (VAEs) face a trade-off between generative quality and coherence-i.e., they struggle to generate realistic and diverse samples that, at the same time, are semantically consistent across modalities. A recent work shows that using a simple approximation to Hölder pooling as an aggregation method improves coherence over the SOTA MMVAE+, despite assuming a single shared representation across all modalities. Yet, it slightly compromises sample diversity. Inspired by this insight, we propose Hölder++, a novel multimodal VAE that improves t...
Computational creativity in Interactive Fiction faces a fundamental tension: Large Language Models (LLM) may produce creative narratives but struggle with world coherence, while symbolic systems ensure consistency but lack creative flexibility. We present IVIE (Incremental & Validated Interactive Experiences), a neuro-symbolic approach to generating complete and playable interactive fiction worlds from scratch. Building upon PAYADOR's neuro-symbolic framework, IVIE implements a four-stage incremental generation pipeline that delegates creative decisions--setting and character creation, puzzle...
Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models (LLMs). However, existing RLVR methods often encourage unnecessarily long reasoning rollouts, which can degrade reasoning coherence and exhaust the available context budget. Existing approaches to long-context organization often depend on external mechanisms to organize rollouts, rather than enabling the model to manage its own reasoning trajectory. To address this limitation, we propose ReSum, a novel RLVR framework that enables LLMs to compress and organ...
Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approaches treat this gap as a knowledge deficit, relying on supervised fine-tuning (SFT) to ingest labeled spatial data from external vision sources or synthetic engines. In contrast, we argue that for many tasks, spatial reasoning capabilities are already present in pre-trained LRMs but require alignment through logical coherence under geometric 2D and 3D constraints. In this work, we propose a self-supervised reinforcement learning (RL) framework...
Introducing North Mini Code: Cohere's first open-source agentic coding model. Built for sovereign developers, this efficient 30B MoE model delivers strong software development performance with minimal hardware requirements.
Generating coherent and controllable long-form content remains a persistent challenge for Large Language Models (LLMs). While reasoning-enhanced models have demonstrated success in logic-intensive domains, our evaluation reveals that they suffer from a severe length collapse in open-ended writing, where performance degrades sharply as target lengths exceed 2,000 words. We attribute this failure to the limitation of static hierarchical planning, which struggles to provide dynamic guidance over extended contexts. To bridge this gap, we introduce the Interleaved Structural Chain-of-Thought (IS-C...
Whisper, a widely adopted ASR model, is known to suffer from hallucinations - coherent transcriptions generated for non-speech audio entirely disconnected from the input. We investigate whether hallucinations can be detected and mitigated through Whisper's internal representations. We extract audio encoder activations and evaluate two representation spaces: raw Whisper activations and Sparse AutoEncoder (SAE) latents. We show that both spaces encode linearly separable hallucination-related information, with discriminative power concentrated in a sparse feature subset and increasing toward dee...
Indoor scene generation is crucial for robot simulation and modern interior design. However, complex layouts together with scarce 3D scene data make learning-based generation challenging. Existing methods often rely on hand-crafted rules or focus on isolated sub-tasks (e.g., floorplan synthesis or single-room furnishing), producing whole-home scenes that lack global coherence, realism, and simulation readiness. To mitigate these limitations, we propose a unified hierarchical framework that decomposes indoor scene synthesis into controllable stages. First, we curate a large-scale dataset of 30...
A blog about how Cohere Labs built coplot, a data visualization tool that not only helps their releases, but also their research process.
Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly observable. Many such problems require imaginative perception: inferring what would be seen from an unseen viewpoint, tracing paths through occluded spaces, or integrating partial observations into a coherent spatial representation. We introduce Imaginative Perception Tokens (IPT), intermediate perceptual representations that externalize what a VLM would perceive under alternative spatial configurations while remaining consistent with the observed input. To stu...
Conversational AI agents require memory systems that are both scalable and semantically coherent across long interaction horizons. Existing approaches rely predominantly on large language model (LLM)-based summarisation at write time, which introduces non-determinism, escalating token costs, and opacity in pruning decisions. We present the Deterministic Memory Framework (DMF), a CPU-first approach that replaces generative memory compression with a fully deterministic pipeline grounded in classical NLP analysis, vector geometry, and mathematical scoring. DMF assigns each conversational interac...
Effective "all-team" summarization in high-complexity settings like the Neonatal Intensive Care Unit (NICU) requires aggregating insights from diverse disciplines (physicians, nurses, therapists) spread across hundreds of clinical free-text notes. Simply pooling heterogeneous text often leads to incoherent outputs. Structured summarization therefore first requires accurate categorization of sentence-level provenance across multi-source notes. This pilot study introduces a clinical provenance categorization pipeline using supervised fine-tuning (SFT) of large language models (LLMs). We adapted...
Order-agnostic language models (OALMs), including discrete diffusion language models (dLLMs), are trained to predict masked tokens under arbitrary conditioning sets, allowing sequences to be generated or scored under arbitrary reveal orders at inference time. In LLaDA-2.1, we report three findings. First, the learned conditionals are not exact factorizations of a coherent joint distribution: changing only the reveal order shifts target log-likelihood by up to 0.49 nats/token, so likelihood alone mixes content difficulty with path-dependent artifacts. Second, although confidence-first (CF) dec...
Unsupervised skill discovery (USD) aims to learn diverse behaviors without reward functions, but often results in task-irrelevant or hazardous behaviors due to uniform exploration. Guided skill discovery (GSD) addresses this issue by incorporating human intent to focus exploration on meaningful regions. However, existing GSD methods typically require training additional guidance models, and rely on pre-defined rules or expert demonstration, which can be ineffective under sparse, online-collected human feedback. To overcome this, we propose COLLIE, a GSD framework that leverages dense unsuperv...
Multi-component LLM agents assemble probabilistic claims from components that each see only part of a joint problem; the composition can violate basic probability axioms even when every component is locally coherent. We formalise this locally coherent, globally incoherent failure via the compositional residual eps*, the L2 distance from the composed quote to the joint coherent polytope, computable at runtime from system output and the declared cross-component coupling constraints. A product-structure dichotomy characterises when local coherence suffices, and a Rayleigh-quotient prediction mat...
A specialized translation model leverages RWS’ global language and cultural expertise and Cohere’s Command A+ model to power the new Language Weaver Pro.
Delivers advanced reasoning with a minimal compute footprint. Command A+ offers full data sovereignty for governments and regulated industries worldwide.
Cohere and Mila announced plans for a new academic research collaboration focused on improving AI evaluation across languages and cultures, starting with French-language cultural context in Quebec.
Recent generative models have largely closed the gap on low-level artifacts - pixel fingerprints, frequency anomalies, upsampling traces - particularly in person-centric and partial-edit settings where the manipulated region is small and surrounded by photometrically authentic content. We introduce Social Gaze Consistency, a high-level semantic cue defined as the mutual coherence of gaze direction, head-eye alignment, and pupil placement between interacting individuals, and show that it constitutes a previously underutilized detection axis orthogonal to existing low-level paradigms. We instan...
Browse the new Cohere Merch Store and view the latest collection for sale, along with an archive of our merch and swag history.
Developer fine-tuned Cohere Transcribe to add diarization and timestamp support, extending open-source speech-to-text capabilities.
Cohere launches Command A+, first open-weights MoE model emphasizing efficiency and latency over peak performance.
Cohere releases Command-A-Plus-05-2026 bfloat16 model weights on Hugging Face Hub.
Cohere signs MOUs with Indra Group and Multiverse Computing to advance AI deployment with focus on sovereignty, security, and accessibility.
Cohere releases Command A+, an open-source model optimized for enterprise agent deployment with improved speed and capability.
Cohere acquires Reliant AI to strengthen sovereign AI capabilities for regulated healthcare and life sciences sectors.
PDI-Bench: Quantitative framework for auditing geometric coherence in generated video via perspective distortion and point-tracking metrics.
Value-filtered decoding selectively applies safety steering at test-time, avoiding unnecessary interventions that degrade helpfulness and coherence.
Cohere outlines governance frameworks for responsible AI development and deployment.
Cohere showcases AI agents for financial services compliance, efficiency, and customer trust.
Cohere argues private AI deployment offers banks control, security, and customization advantages.
Cohere advises financial firms on scaling generative AI from pilot programs to production with risk management.
Cohere positions LLMs as solutions for financial services knowledge work and customer service automation.
Cohere discusses metrics and integration strategies for achieving AI-native enterprise operations.
Proposes segment-level supervision for LLM-based Lean 4 theorem proving, balancing dense local signals of step-level training with coherence of whole-proof generation.
Google's leaked video model 'Omni' shows improved text coherence in generated video content.
Cohere publishes enterprise AI governance guide covering monitoring, responsibility frameworks, and innovation-risk balance.