Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
30B MoE analysis: pretraining loss alone fails to predict post-SFT performance; solution density better predicts downstream quality.
Every story tagged with this topic, ordered by date.
30B MoE analysis: pretraining loss alone fails to predict post-SFT performance; solution density better predicts downstream quality.
Ostrich: GPU-accelerated differentiable rigid-body simulator optimizing contact gradients via tape compression; enables gradient-based robotics optimization.
Cohere releases North Mini, a decode megakernel LLM serving engine achieving 1.58× speedup over vLLM in production.
Linux kernel git infrastructure struggles with scraper traffic consuming more resources than legitimate access, raising infrastructure cost concerns.
Analysis of Astra and frontier models' Architecture/Engineering/Operations (AEO) choices, tracking patterns relevant to founders and developer experience leaders.
DFlow enables verifier representation reuse in block diffusion speculative decoding, improving LLM inference by recovering computation from rejected tokens.
Study characterizes interaction between inference-time activation steering and weight-only quantization (INT8, NF4) on 7-9B open-weight LLMs.
Hierarchical Wasserstein Merging addresses multi-domain multi-task learning by merging specialist model distributions rather than parameters, reducing interference under distribution shifts.
Linear algebra unification of 14 works on transformer attention rank, compression, and KV-cache efficiency via SVD and eigendecomposition.
Parameter-efficient LoRA adapters for MRI reconstruction across variable acceleration factors via shared weights and gating.
Theoretical bounds on query-oblivious coresets for softmax attention; improved gap closure between upper and lower complexity limits.
Adaptive early-exit system for report generation predicting per-section computation value; 4x latency reduction under low budget.
LSTM-based sequence prediction detects sensor impersonation in unauthenticated IoT deployments via anomalous temperature readings.
SeaCausal-FL federated framework for maritime engine fault diagnosis combines interval type-2 fuzzy logic with mechanism-conditioned causal models.
RegionFed: personalized federated learning for retail search across heterogeneous regions, fixing transformer collapse in parameter-level FL.
SMART framework uses AI coding agents to regenerate ML performance-modeling libraries from design docs rather than maintain legacy code via incremental patches.
Layer dropout re-evaluated for modern LLM pre-training; comprehensive study establishes conditions for improved efficiency without accuracy degradation.
Proton irradiation testing of open-source Tensil NN accelerator on Zynq platform for radiation-hardened spaceborne computing.
FedDRAW applies federated learning with reputation-weighted aggregation to multi-institutional chest radiograph classification.
DEX-Comp: two-stage training recipe for soft context compression in RAG systems via pure distillation and fine-tuning.
arXiv proposes time-to-first-spike SNNs with reference-based encoding for energy-efficient large language models.
NS-ST-GraphRAG: neuro-symbolic framework for literary narratives integrating spatio-temporal constraints and ontology-guided extraction for RAG over long-form text.
SparseBin: tuned GPU self-organizing map algorithm achieving 5.6-10.1x speedup on best-matching-unit search via tile optimization and vectorization.
Neural function compilation via teacher-student training converts natural-language specs into compact, reusable adapters, reducing inference cost and latency vs. remote model calls.
Para-Pipe optimizes neural network latency on heterogeneous SoCs via hierarchical operator parallelism for edge inference.
Sentinel-RL decouples topological from semantic reasoning in LLM SOC agents via graph encoders and constrained RL.
Ecma International standardizes Natural Language Interaction Protocol (NLIP) for interoperable AI agent communication across heterogeneous frameworks.
Direct-P algorithm optimizes FP4 attention on NVIDIA Blackwell via direct quantization path, achieving 2.13× BF16 throughput for causal inference.
NVFP4 W4A4 quantization on all 496 linear layers of Qwen3.8-27B hybrid LLM including Gated DeltaNet recurrent blocks maintains performance across evals.
TAP-Path compresses Virchow2 pathology encoder via task-adaptive block selection and token pruning for inference efficiency.
Diffusion-augmented LLMs enable parallel token generation from autoregressive distributions, decoupling NTP and diffusion weights for lossless speedup.
llm-anthropic 0.28 displays Claude reasoning traces by default and adds refusal exception handling.
Graph Machine architecture uses dynamic sparse routing with O(n) state for efficient pretraining; reduces Qwen3-0.6B compute by replacing 75% dense layers.
JAX library for computing reverse-mode gradients of ODE ensembles at native solver speed, targeting scientific and engineering applications.
Proposes FP4 pretraining recipe with E5M3 block scaling and selective stochastic rounding to stabilize 4-bit language model training.
llm-gemini 0.34 adds Gemini 3.8 Flash support with configurable thinking levels and async bug fixes.
ShallowStream optimizes streaming video MLLMs by pruning computation at shallow layers, reducing overhead for embodied AI and autonomous driving.
Factory edge agents selected via retrieval-augmented answer quality outperform parameter-count heuristics for on-premise deployment.
LoRA-TSD optimizer treats LoRA updates as tangent vectors on fixed-rank manifolds, achieving 2.8× speedup.
H3DNAS compresses 3D point cloud models directly on ONNX binaries for edge deployment without source code access, targeting NVIDIA Jetson Orin Nano.
Multi-shell Leech-lattice decoder implementation for 2-bit LLM quantization with GPU kernels optimized for serving-phase GEMV at batch 1.
Differentiable optimization layer enables gradient-based planning for data center electricity costs by backpropagating through market-clearing constraints.
Debias-SparseGPT addresses bias amplification in pruned LLMs using representational debiasing with demographically contrasting inputs.
ViSAR uses training-free adaptive retrieval for Document VQA, dynamically selecting page count per query to reduce LVLM latency.
Scalable Kronecker-Fisher approximation enables Hessian analysis for billion-parameter LLM compression without full Fisher matrix storage.
Systems survey of GUI agents analyzes observation, memory, action, and runtime efficiency across web, mobile, and desktop environments.
Anthropic releases Claude Fable/Mythos 5.1 with SOTA performance, 75% cache cost reduction, and 70% increased output token throughput.
Causal analysis across 9 open-weight models shows quantization damage is distributed globally, not localized to task circuits or weight statistics.