Multi-Agentic System Leveraging Open-Source LLMs to Mitigate Disinformation Threats
Multi-agent system using open-source LLMs for automated disinformation detection and fact-checking at scale across social media.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Multi-agent system using open-source LLMs for automated disinformation detection and fact-checking at scale across social media.
KnowsTFM: knowledge-graph-informed fine-tuning improves small tabular foundation models in niche, high-dimensional, low-data domains.
EMPATH: multilingual multi-turn safety benchmark for emotional-support chatbots across 140 scenarios, 34 personas, and 19 metrics on crisis handling.
Import AI 463 covers self-improving robots, a 10k-GPU Chinese cluster deployment, and commentary on AI's societal transition.
Inoculation adapters: LoRAs that suppress undesired capability generalization via three-stage training and disposal, reducing emergent misalignment backdoors.
Omen AI raised a $31 million Series A to monitor chip coolant and stop bacterial outbreaks in data centers.
Unsupervised community detection in heterophilic graphs using discrete Ricci curvature and diffusion.
Text-to-video diffusion models leverage data manifold structure as implicit reward signal without external annotations.
Theoretical framework for grounding LLM reasoning with incomplete knowledge graph evidence using entity anchors and path energies.
Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated assistant task or a closed workflow. Therefore, autonomous science needs a collaboration infrastructure that coordinates projects, agents, and digital and physical resources. We identify this as a shift from code-centered execution loops to research-oriented collaboration processes, where questions, evidence, participants, and resources must be coordinated under uncertainty. In this framing, an agent may be an AI system, a human researcher, a team, a labor...
In our goal to develop personalised dysarthric speech recognition (DSR) models, this study compared the recognition performances of human listeners and those of three state-of-the-art, off-the-shelf ASR systems (Whisper-large-V3, Google Chirp 3, and Omnilingual) on the recognition of Dutch continuous read and spontaneous speech from a single speaker with severe dysarthria. Results showed that both humans listeners and the three off-the-shelf ASR systems exhibit word error rates (WER) exceeding 70% on average, indicating that DSR is highly challenging for both humans and ASR systems. Fine-tuni...
Medication errors, particularly dosing errors in clinical trials (CT), can lead to patient harm, adverse drug events and worse patient outcomes. Dosing errors are preventable, and early identification can improve trial integrity and mitigate subsequent clinical and financial burden. This study aims to detect dosing errors within CT protocols by evaluating text representations of trial information using transformer-based language models trained on biomedical corpora. CT textual data was encoded using several models, including ClinicalBERT, PubMedBERT, BioBERT, and MedCPT, and integrated with c...
Learned reconstruction operators for inverse problems are typically trained under a fixed noise model, and generalize poorly when the distribution during testing differs from the one assumed during training. Distributionally robust optimization (DRO) addresses this by optimizing against the worst-case distribution within a prescribed ambiguity set, but standard Wasserstein DRO perturbs the full joint distribution uniformly, which can be overly conservative and ignores the physics of the measurement process. We develop a structured DRO framework in which the ambiguity set is restricted to stru...
Modern engineering workflows increasingly rely on massive parallel simulation, driving the need for scalable, large-batch Bayesian Optimization (BO). Existing batch BO methods, however, incur large computational cost or rely on approximations that erode batch diversity. We propose B3O (Boltzmann Batch Bayesian Optimization), a framework that reframes batch generation as a pure sampling problem: drawing samples directly from the Boltzmann distribution defined by the acquisition function avoids the bottlenecks of existing large-batch methods. Theoretically, we prove that queries sampled from th...
Hessian spectral properties are a standard tool in analysing neural-network training, with eigenvalues linked to sharpness, generalization, and optimization dynamics. Eigenvalues quantify curvature magnitude, while eigenvectors identify which parameters generate that curvature. In this work, we study how the leading Hessian eigenvectors evolve during training and how they affect the learning trajectories. We track the training dynamics of multilayer perceptrons on a classification problem and measure eigenvector dynamics through two complementary statistics: (i) displacement over time, inspir...
LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluation, LLM-as-judge reliability, safety evaluation, jailbreak/refusal robustness, reward hacking, mecha...
Proposes adaptive routing between draft and target multimodal models for efficient visual reasoning by learning query difficulty signals.
Introduces sparse cross-modality fusion for RGB-T object detection, reducing computational cost via selective thermal-visible integration.
Multi-center breast cytology dataset with 470 whole-slide images annotated for C1-C5 classification in pathology AI.
Philosophical essay framing data centers as the embodied substrate of modern AI systems using biological analogies.
SHOVIR benchmark evaluates whether radiology report generation models rely on image evidence or learned priors via spatially-grounded metrics.
Analyzes SONAR multimodal embeddings to detect encoding anomalies via dimension sensitivity analysis.
Sim-to-real transfer remains a major obstacle for reinforcement learning (RL), especially for vision-based control where image observations exacerbate the state-distribution shift between simulation and the real world. Domain adaptation (DA) is a promising remedy for this challenge. Prior sim-to-real DA works have demonstrated encouraging results, yet these approaches typically assume substantially more target data, which is not available in practice. Indeed, their performance degrades significantly when the target data budget is reduced. To address this challenge, we propose AIDA (Adaptive I...
How does an agent that can tell self from world come to be durably shaped by that distinction? Recent work shows that a predictive system can detect its own agency (Ye, 2026), but detecting agency does not explain durable, self-shaped behavior. We show that agency-gated slow credit -- a conjunctive term Own*Agency*Salience driving a slow parameter update -- produces post-unload behavioral residue: on a spiking substrate (Nengo LIF/PES), a learned self-preserving choice survives episodic buffer removal (retained fraction 0.96, N=50) and collapses when the slow decoders are reset or the agency ...
Existing domain-incremental learning (DIL) strategies call for massive amounts of data to adapt to new domains and suffer from the overfitting problem in the case of data scarcity. This paper puts forward a relatively uncharted problem, namely, few-shot domain incremental learning (FSDIL), taking into account the problem of extreme data shortages in the realm of DIL. A novel algorithm, namely Continual Vision-Language Consolidation (CVLC), is proposed to address the FSDIL problem, where the key idea lies in the concept of latent space reservation in the base domain coupled with dual coalescen...
Current multimodal fusion approaches, particularly those based on static Mixture-of-Experts (MoE) architectures, often struggle to provide the adaptive and efficient collaborative reasoning required by complex real-world applications. We introduce the Dynamic Agent-based Interaction Network (DAIN), which reconceptualizes multimodal fusion as a dynamic, multi-agent collaborative process. DAIN employs a context-aware Meta-Controller that dynamically schedules sparse activation of specialized interaction agents and orchestrates compressed inter-agent communication for consensus-building. The fra...
Dynamo enables frozen VLMs to evolve reusable skills and visual tools without retraining, improving reasoning on visual tasks.
MirrorCode benchmark evaluates AI coding agents on full program reimplementation tasks, extending beyond short coding exercises.
Cortex organizes web-scale training corpora via ontological graphs instead of flat filtering, improving LLM data pipeline structure.
NMO Benchmark evaluates generative molecular design on quantum materials science tasks, beyond traditional drug discovery benchmarks.