Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark
24 frozen vision foundation models (ViT, CLIP, etc.) evaluated for face presentation attack detection via linear probing under domain shift.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
24 frozen vision foundation models (ViT, CLIP, etc.) evaluated for face presentation attack detection via linear probing under domain shift.
Algebraic framework deriving transformer expressivity from finite-precision attention dynamics and memory constraints.
The startup analyzes calls, messages and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.
SymmGrid framework accelerates on-robot reinforcement learning via parallelized symmetries and dual visual perception.
LLM-based validation of atom-centered structural descriptors reveals descriptor degeneracies in materials science.
OptimismBench detects directional bias in LLM probability judgments via inverted-pair evaluation method.
TREK benchmark evaluates LLM agents on complex, executable travel itinerary planning with verifiable constraints.
Pair-level judgment outperforms dialogue-level generation for emotion-cause extraction in conversation datasets.
Feature stability analysis framework shows feature bagging ensemble reduces generalization error.
BAND sparse Bayesian network achieves polynomial convergence rates for high-dimensional distribution estimation.
CreditCardQA benchmark evaluates LLM numerical reasoning on real financial documents with CoT/PoT comparison.
Progressive Multimodal Alignment mitigates projector drift in continual instruction tuning of vision-language models.
Belief-Guided architecture reduces Monte Carlo Tree Search dependency in Computer Go, improving inference on consumer hardware without deep search hallucinations.
Surrogate-guided dual-objective search method for neural ensemble search addresses computational intractability of joint architecture and composition optimization.
Three-class detection framework distinguishes humans, bots, and AI agents in browser automation traffic; binary classifiers misclassify 39.1% of agents as human.
FedDAB defends federated learning against backdoor attacks via model-contrastive regularization and alignment checking, addressing statistical heterogeneity.
Paired-prompt methodology tests whether language models correctly match diagnostic evidence to causal claims across different populations and estimands.
Latent-IM recovers state estimation and action control for conversational moves in speech LLMs by decomposing move selection and realization in hidden representations.
Temporally centered SIGReg improves multi-task latent world-model learning by maintaining task-dependent cluster separation during marginal Gaussianization.
Two-call self-refinement outperforms five-agent pipeline on Qwen2.5-7B; multi-agent systems suffer error accumulation, dropping GSM8K accuracy to 45% with JSON format.
BioVLN simulation platform for visual-language navigation in biomedical labs with instrument-specific approach constraints and safe clearance requirements.
DuPLeR dual-path LLM reasoning framework combines multimodal and LLM priors for few-shot knowledge graph completion while filtering hallucinated evidence.
Formalizes detection of outcome performativity—when predictions causally influence predicted outcomes—via A/B intervention testing in palliative care, credit, recommender systems.
Pegasus framework translates human manipulation videos into robot-learnable data via task graphs and affordance constraints, addressing embodiment gap for embodied AI.
Preregistered experiment: diverse human groups (L1/L2 English writers) produce more collective creativity than LLMs; AI assistance risks homogenizing or replacing human diversity.
AI is fantastic at spotting patterns, but human insight is the key.
DIRECT framework applies DPO and controlled decoding to improve LLM sequence labeling alignment and inference efficiency for information extraction tasks.
CoRAS adaptively selects acquisition/compression rates for high-resolution imaging using conformal prediction to bound reconstruction error probabilistically.
Claude Opus-4.7, GPT-5.4, Gemini-3.1-Pro confabulate medical diagnoses without images; diagnosis systematically shifts by patient demographic, raising safety concerns.
SERPO enables test-time LLM self-improvement via co-evolving rubrics and evidence for open-ended generation without external reward models.