Evidence-Bounded Mental Health Reasoning from Heterogeneous Speech Protocols
Framework addresses evidence-bounded mental health reasoning from heterogeneous speech protocols, preventing hallucination across different clinical tasks.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Framework addresses evidence-bounded mental health reasoning from heterogeneous speech protocols, preventing hallucination across different clinical tasks.
TSPFN adapts TabPFN foundation model architecture for physiological time-series classification with temporal dependencies in low-data regimes.
LiFT uses language-informed Flow Matching with agent-guided SMILES generation for 3D molecular design balancing affinity and chemical validity.
Clinical NLP framework uses automatic language analysis to augment psychiatrist judgments of patient experience in therapeutic conversations.
Study finds KV-cache quantization in RAG systems reduces faithfulness to retrieved evidence even when accuracy holds on Qwen2.5-7B.
Knowledge-aligned SFT framework constrains fine-tuning targets to base model's parametric knowledge to reduce hallucinations via Evidence Rewrite and Recall variants.
Stiefel-constrained Riemannian optimization enables parameter-efficient rotational activation steering to control LLM refusal behavior at inference.
New benchmark and synthetic training pipeline improve LLMs' self-modeling—ability to predict own behavior on counterfactual prompt edits.
Sparse competition-based training induces modular specialization in neural networks while maintaining baseline accuracy.
CastClaw: human-in-the-loop LLM agent system for industrial time-series forecasting with domain-specific constraints and uncertainty communication.
MR-JEPA: self-supervised video foundation model for cardiac MRI pretraining on multi-sequence data from 10,505 patients using 3D spatiotemporal masking.
CoJEPA combines JEPA and contrastive learning for music representation learning, balancing local task performance with global representation strength.
Debian voted to allow developers to use AI tools in their contributions to the Linux distribution's "development, maintenance, [and] documentation." The new policy on AI acknowledges that "responsible" use of AI can improve developers' productivity, and goes on to say, "generative AI is neither exempt from nor subject to special rules beyond the standards already expected of Debian contributors." The project's voting developers considered several other proposals on how to handle AI, including some that would have banned contributions made using AI tools. As It's FOSS reports, some users and c...
Speculative philosophical narrative about AI agents forming emergent social structures by 2026, invoking social contract theory.
Nvidia invests $3.5 billion into Taiwanese chipmaker MediaTek. The deal shows how Nvidia plans to stay essential to AI infrastructure as Big Tech begins to build its own AI chips.
Anthropic outlines updates to alignment and security research, internal testing processes, and risk mitigation strategies for AI systems.
Today, I’m talking with New York Governor Kathy Hochul, and I’ll just warn you — this episode moves really fast. It’s an election year, after all, with a shocking amount of tech policy at stake, and Governor Hochul has taken strong positions on almost every major tech issue there is. For example, Meta just reached a settlement with dozens of states, including New York, which will restrict how teens use platforms like Instagram in very specific ways. Governor Hochul is a strong supporter of those restrictions and more, as you’ll hear. But those come with a cost — widespread age verification me...
Explosive growth comes with a new regulatory burden in the European Union.
Import AI newsletter covers Hugging Face concerns, space mining applications, and Five Eyes AI governance positions.
OpenAI will soon be held accountable for mitigating risks related to ChatGPT's impact on minors, user mental health, and the spread of illegal content in the European Union. That's because ChatGPT is now considered a Very Large Online Search Engine under the EU's Digital Services Act, a set of laws regulating major online services and platforms. Along with ChatGPT, the European Commission also announced that Reddit and Roblox are considered Very Large Online Platforms, subjecting them to the same rules. The DSA restricts platforms from targeting ads to minors or using a person's sexual orient...
“AI creator” accounts like Aitana Lopez will get a new “AI-generated profile” label. | Image: Aitana Lopez Instagram is finally taking steps to address the rise of fake AI-influencer accounts that have gotten harder to spot. It's also renaming the "AI creator" label to "AI-generated profile" to make it clear when a profile features an AI-generated person that's not a real human being. "We've heard that people don't like seeing a profile that seems human, only to find out later that the person featured is AI-generated," said Instagram in a blog post. "They want to know when a profile features ...
Circleback is also introducing new pricing plans starting from $14 per month
Murano: open-source framework for composable mechanistic interpretability pipelines across loading, attribution, intervention, and evaluation stages.
SwarmBench: benchmark for evaluating LLM-based agent orchestration across accuracy, efficiency, cost, and process quality in dynamic multi-agent systems.
Hybrid CNN-Transformer for Antarctic sea ice forecasting using seasonal factorization; domain-specific climate modeling, not AI capability research.
MIOH: fine-grained multi-image object hallucination benchmark for MLLMs, diagnosing how visual complexity triggers factual errors across multi-image reasoning.
PyKEEN-NSX: modular negative sampling framework for knowledge graph embedding; specialized infrastructure for KG link prediction, not core LLM research.
Geometry of Divergence: framework tracking LLM hidden-state trajectories via temporal curvature and variance to diagnose representation drift in multi-turn reasoning.
Retain-free identity unlearning in MLLMs via localized early-decoder MLP intervention, enabling post-deployment privacy removal without retain sets.
Empirical analysis of precomputed memory (KV caches, compressions) degradation in Llama-3.1-8B: assembly failures and rebuild costs nearing full computation.