When Persona Attributes Improve Population Alignment in Large Language Models
Persona prompting steers LLMs toward survey-participant response alignment; empirical study of attribute effectiveness for population prediction.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Persona prompting steers LLMs toward survey-participant response alignment; empirical study of attribute effectiveness for population prediction.
QUEST improves uncertain knowledge graph completion via spectral initialization and scheduled smoothing without adding trainable parameters.
Orthogonal ensemble of 11 models achieves 36.8% macro-F1 on leave-performer-out emotion recognition from skeleton motion, +43% over baseline.
Challenges exponential-moving-average teacher-student design for test-time adaptation; shows error accumulation persists on longer sequences.
Edelson PC is filing 30 new lawsuits against OpenAI over the Tumbler Ridge shooting, escalating claims to aiding and abetting and naming Chris Lehane, though evidence remains unconfirmed.
VMetaphor-Bench: first benchmark for visual metaphor generation in text-to-image models; 1,500 curated samples across three levels and ten categories.
Empirical study shows single random seed in recommender-system experiments is risky; seed choice affects hyperparameter selection and final conclusions.
RINSE: gradient-free zero-shot graph anomaly detection via target-time normality self-estimation, handling domain shift without retraining.
Debias-SparseGPT addresses bias amplification in pruned LLMs using representational debiasing with demographically contrasting inputs.
ATV Big Air Tour reduced marketing and merchandising tasks from 3 days to 3 hours using ChatGPT Work, including auto-generated inventory website.
ViSAR uses training-free adaptive retrieval for Document VQA, dynamically selecting page count per query to reduce LVLM latency.
Comparative analysis of worldbuilding strategies in AI-generated vs. human fiction across GPT-4, LLaMA 3.3, Mistral 3.2, Gemma 3.
PragAlign framework for controlled synthetic dialogue generation using LLM-based evaluators in generate-evaluate-revise loops.
Behavior-based fusion model combining LLM differential diagnoses with ontology rankers for rare-disease diagnosis interpretability.
DeepAffinity applies Small Language Models to predict eCommerce user product preferences (brand, size, color) via temporal prediction.
CivBench benchmark for long-horizon LLM agents in Civilization VI via Model Context Protocol, spanning 300+ turns and 76 tools.
ICE-T didactic framework addresses ML education opacity via intermodal transfer and computational thinking for calibrated AI trust.
Scalable Kronecker-Fisher approximation enables Hessian analysis for billion-parameter LLM compression without full Fisher matrix storage.
CACTUS enables semantic backdoor attacks in decentralized federated learning using mask-guided, modality-specific trigger embedding.
Threat Conditional Network (TCN) enables single model robustness across diverse adversarial threat levels via representation factorization.
Open-weight transformers show logical validity representations remain decodable from hidden states despite near-chance behavioral performance.
IFW-BLS introduces intuitionistic fuzzy loss for robust broad learning systems handling noisy labels and outliers.
UTP-Bench evaluates LLM travel planning under real-world uncertainty including delays and stochastic disruptions.
Multi-turn agent credit assignment succeeds under high verifier information density; coverage strategy beats targeting in low-density regimes.
BLUEPRINT jailbreak framework uses Monte Carlo Tree Search over 18 theory-grounded influence factors to achieve near-ceiling attack success on frontier models.
Sparse MoE routing decisions across layers share common geometric structure aligned via control subspaces despite layer-specific coordinate systems.
Contrastive explanations for Quantitative Bipolar Argumentation Frameworks explain differences between arguments rather than single outcomes.
Retrieval-augmented generation and BioNER improve lay summarization of radiology reports to reduce patient reliance on general LLMs.
PolERo dataset evaluates political evasion classification in Romanian, testing cross-lingual transfer of English response-clarity taxonomy.
Fable 5.1 removes data retention policy and increases caching; Stratechery commentary on enterprise safeguards.