PeTeR: Post-Training Robustification of Probabilistic Circuits
PeTeR applies post-training Wasserstein-robust optimization to probabilistic circuits for improved generalization under distribution shift.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
PeTeR applies post-training Wasserstein-robust optimization to probabilistic circuits for improved generalization under distribution shift.
Analysis of Bielik LLMs (1.5B-11B) showing activation dispersion in MLP layers predicts entity familiarity and factual reliability across domains.
DiaLLM studies English dialect generation in open-weight LLMs via continual pretraining on ICE, comparing alignment strategies across Australian, Indian, British English.
Analysis of classifier-free guidance overstabilization in diffusion samplers with proposed terminal-fitted repair for high-guidance regimes.
Survey of 1,250 arXiv papers (2024-2026) categorizing recursive self-improvement in AI systems by loop closure and improvement target.
Meta is adding a new safeguard to stop people from secretly recording others with its AI glasses. But the update comes as the company continues to expand how much personal data its AI products collect and use.
RL post-training composes primitive skills into novel reasoning strategies beyond base model capabilities in controlled symbolic environments.
OpenAI is overhauling ChatGPT's voice mode with a new model that it says is more like "talking to another person." The new GPT-Live-1 is designed to interrupt you less and will also wait for you to continue speaking if you pause mid-conversation. During a press briefing, OpenAI research lead Kundan Kumar called GPT-Live-1 the company's "smartest voice model" yet. It will automatically pass your queries to its best text models, like GPT-5.5, when it needs to reason or search the web, allowing it to more quickly transition from researching the topic you've asked about to talking about its findi...
OpenAI says its new voice mode can speak and listen at the same time, a key ability for live translation.
ALER-TI retrieval-augmented framework imputes time series using historical patterns to handle non-stationary dynamics.
Optimal control framework adapts neural network depth via posteriori error estimation for nonlinear problem approximation.
Hybrid quantum-classical architecture combines quantum neural networks with path signature kernels for time series classification.
Future Confidence Distillation measures pre- and post-solution confidence in LLM responses to improve reliability estimation for downstream tasks.
Android Bench is evolving, and developers can help guide that process.
Riemann normal coordinates improve Levenberg-Marquardt optimization for nonlinear least-squares via geodesic acceleration.
Systematic review of agentic AI governance identifies governance priorities and mechanisms for autonomous planning/execution systems.
Asymmetric focal loss improves graph neural network prediction of polypharmacy drug-drug interactions via difficult-example emphasis.
The round, led by Radical Ventures, values the two-year old startup at $1 billion.
CARLA-GS decouples representation, reasoning, and physics for photorealistic autonomous driving corner-case synthesis.
Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance... Presto is an open source, distributed SQL engine for running fast, interactive queries on very large datasets. On NVIDIA GPUs, Presto delivers peak performance for analytical query workloads and provides low latency for users and agents. GPU-accelerated Presto brings low latency to your analytical workloads, keeping you and your agents unblocked and iterating as fast as possible. Source
Multi-label BERT formulation outperforms multi-class for CVE-to-CWE vulnerability classification across taxonomy sizes.
FedKT-C combines one-shot federated learning with synthetic data generation to achieve privacy, low communication, and heterogeneity robustness simultaneously.
PALS adjusts per-layer sparsity in LLM pruning via activation percentiles, improving LLaMA-2-7B perplexity by 15% at 50% sparsity over uniform Wanda.
Empirical study on capacity allocation across hierarchical search agent roles (delegation, execution, generation) in multi-agent LLM systems.
There are a lot of fast-growing AI startups, but some are growing even faster, they say.
Bounds on failure region probability for Langevin dynamics training, proving exponential suppression of unsafe states with dimension.
Mahalanobis distance-based unified detection framework for LLM-generated text, hallucinations, watermarks, and adversarial examples.
Analysis of human-AI collaboration in structural design showing friction from constraints can drive creative exploration rather than final-answer generation.
GRiLS, a gradient-free Riemannian MCMC sampler for multimodal distributions, improves mode mixing without density gradient evaluations.