TSAI-MetaFraud: A Benchmark Dataset for Financial Fraud Transaction and Behavioral Risk Detection in Metaverse Ecosystems
TSAI-MetaFraud: multimodal benchmark dataset for fraud and risk detection in metaverse virtual economies.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
TSAI-MetaFraud: multimodal benchmark dataset for fraud and risk detection in metaverse virtual economies.
ALICE: unified pathology foundation model via multi-stage distillation from eight expert models, trained on 25M+ pathology images.
SAGEAgent learns to selectively acquire multimodal clinical data for cancer survival prediction, reasoning about cost-benefit tradeoffs in diagnostic workflows.
Energy profiling of edge VLM inference reveals language generation dominates energy costs, not vision processing, across five models and two hardware platforms.
Large-scale study of LLM CLI coding-agent failure trajectories reveals onset patterns and recovery mechanisms, treating failure as temporal process not final outcome.
VGGT geometric foundation model encodes co-visibility of image pairs as emergent 3D-aware representation, useful for reconstruction and robotic localization.
Rashomon Explanation paradigm couples prediction with explainability as complementary objectives, showing self-explanation can improve rather than degrade model accuracy.
Square Root and Hapax Correction techniques derive well-calibrated likelihood ratios for forensic authorship verification without auxiliary calibration data.
Shared Selective Persistent Memory architecture for agentic LLM systems retains reusable task specifications, schemas, and tool configs across multi-turn sessions.
Multimodal RL reward hacking in MLLMs via visual misalignment; introduces NRFR metric to measure failures in improved-reward samples across VQA and chart tasks.
Terminal embeddings preserve pairwise distances under dimension reduction with applications to k-means and k-median coresets for time-series clustering.
Semantic framework distinguishes AI outputs as engineered representations versus facts, formalizing failure modes like extrapolation, refuted assertion, and stale sources.
Analysis of representational variance in language models shows token-level context dominates (79-91%) over category structure (4-12%) across 14 models, challenging neural collapse theory.
BTHA framework decouples language guidance from vision-text backbones for medical image segmentation via stable feature-level adapters.
Foveated Dynamic Transformer (FDT) applies human visual system principles for robust, efficient vision transformers with inherent noise/adversarial resilience.
ProofCouncil agent using author-critic architecture solved 6/10 real mathematical problems in FirstProof challenge via agentic workflows.
Adaptive Multi-Teacher Routing (ATR) uses uncertainty-driven structure selection to improve universal machine-learning interatomic potentials with limited high-fidelity data.
Pipeline combines Ghidra reverse engineering, anchor-based retrieval, and LLM reasoning to recover source code from stripped binaries.
Test-time prompt adaptation for CLIP exploits distributional brittleness of adversarial perturbations to improve vision-language model robustness without retraining.
Analysis of Bayesian causal discovery under latent confounding in linear Gaussian networks identifies critical correlation thresholds affecting DAG posterior inference.
Parameter-efficient CLIP adapter with continuous metadata conditioning for long-term animal re-identification under morphological and seasonal distribution shifts.
Test-time scaling on EXAMS-V multilingual benchmark shows parseability is primary factor for small VLMs (Qwen2.5-VL-7B, Qwen3.5-4B), not search algorithm.
Open source AI is booming, according to Hugging Face CEO Clem Delangue. The company has grown into something like a GitHub for AI in recent years, where AI builders can share and download open models and datasets, now used by roughly half the Fortune 500. Delangue has seen the same story play out again and again: companies start […]
Multimodal retrieval framework combining vision and trajectory for autonomous-driving scenario matching in large-scale datasets.
Soofi S 30B-A3B: open-source MoE-Mamba hybrid for German/English with 3B active parameters, matches 14-27B dense models on benchmarks.
QADAPT: multi-agent RL with factored action spaces for tuning electrostatically-defined quantum-dot arrays.
SVF-CR: multimodal fusion framework for detecting ambivalence and hesitancy from synchronized facial, visual, and acoustic cues.
Test-time training method using guided self-learning to improve long-context LLM performance on extended input sequences.
Though Instagram head Adam Mosseri doesn't want to filter out AI content on the platform, he argues that you "shouldn't have it in your feed" if you don't like it. "I don't think we should filter out AI content," Mosseri said during an interview on Lenny Rachitsky's podcast. "I think we should let you know if content is AI content or not." At the same time, Mosseri seems to be drawing a distinction between content-based sorting and banning AI from the platform entirely. In fact, he believes people who love AI content "should be able to have a feed that's just AI town." Instagram, like many ot...
Theoretical analysis: InfoNCE contrastive loss population risk is O(1/k)-close to cross-entropy with k negative samples.