Topics
46 curated papers across all published weeks — sorted by disruptiveness score.
Time series reasoning is crucial to decision-making in diverse domains, including finance, energy, and scientific discovery. While existing time series foundat…
Read free explainer →
2026-W34
Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement.…
2026-W35
Anatomically plausible segmentation remains challenging because of low contrast, ambiguous boundaries, and modality-specific artifacts. Interactive segmentatio…
2026-W36
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as su…
2026-W37
Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under conti…
2026-W38
Medical image diagnosis has achieved significant progress with deep learning, yet existing methods often rely on isolated visual evidence and lack the ability…
We present Swift-Image, a compact unified model for text-to-image generation, single-image editing, and multi-image editing. Our goal is to explore how far a r…
We introduce DeepSeek-R1, a reasoning model trained with large-scale reinforcement learning that achieves strong multi-step reasoning without extensive human-a…
2026-W30
Prefix adders are fundamental arithmetic circuits, but their design space grows exponentially with bit-width, posing significant optimization challenges. Previ…
Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing querie…
High-risk non-muscle-invasive bladder cancer (HR-NMIBC) carries substantial risks of recurrence and progression, while current clinical risk stratification rem…
We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstructi…
DeepSeek-V3 is a strong Mixture-of-Experts language model with 671B total parameters and 37B activated per token. We describe Multi-head Latent Attention, an a…
The entire ecosystem of open-source language models effectively relies on a single platform. What if this platform was forced to shut down tomorrow? Implementi…
Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal…
While data-driven 3D shape correspondence estimation has recently seen substantial progress, robust matching under partial observations and strong non-isometri…
Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communication patterns complicat…
Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately.…
We scale sparse autoencoder interpretability techniques to frontier multimodal models, recovering monosemantic features that enable precise activation steering…
We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VA…
We present Llama 3, a family of foundation models that support multilinguality, coding, reasoning, and tool use. We detail pretraining data mixtures, scaling l…
We train a large-scale transformer policy on heterogeneous humanoid datasets that zero-shot adapts to new morphologies and terrains, combining imitation, reinf…
Large Language Models (LLMs) increasingly require selective removal of harmful or sensitive knowledge, called unlearning, yet existing methods and benchmarks f…
Deep neural networks are often overconfident, assigning high confidence even to incorrect predictions. Consequently, users lack a reliable signal for deciding…
Label-free virtual staining offers a compelling, non-destructive alternative to standard histopathology; however, its clinical adoption is hindered by the comp…
Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbolic, auditable, and r…
We prove a rigorous learning-theoretic separation showing that a quantum learner with access to a single clean qubit and mixed states can efficiently learn spa…
Machine-learned drive shaping improves hot-spot symmetry in inertial confinement fusion experiments, yielding a statistically significant boost in neutron yiel…
A multi-satellite fusion pipeline attributes methane plumes to facility-level sources within hours, with quantified uncertainty suitable for regulatory enforce…
An autonomous laboratory combines literature-trained language models, Bayesian optimization, and robotic synthesis to discover three previously unreported soli…
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and…
2026-W33
We integrate high-resolution tactile sensing into vision-language-action models, enabling contact-rich assembly tasks that pure vision policies fail, with stro…
We train neural cellular automata to design biocompatible scaffold growth programs that regenerate complex tissue geometries in silico and guide 3D bioprinting…
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merel…
2026-W31
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis generation and code execu…
We introduce the first Probably Approximately Correct (PAC) learning framework for general-sum concurrent stochastic games (CSGs) with transition uncertainty,…
High-fidelity 3D generation predominantly relies on scaling model capacity and data, which incurs prohibitive computational costs. This paradigm typically requ…
While Multimodal Retrieval-Augmented Generation (MM-RAG) has shown promising results, it still struggles with complex multi-hop reasoning tasks. Existing metho…
Efficiently retrieving relevant clips from large-scale driving logs is essential for data curation, model development, and safety analysis. Structured and rule…
Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic…
Deploying autonomous computer-use agents (CUAs) locally is increasingly important for privacy, cost efficiency, and practical usability, yet improving their pe…
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a…
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source an…
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an a…
Tabular-to-image methods have emerged as novel approaches to leverage the high predictive performance of convolutional neural networks and vision transformers.…
Snapshot Compressive Imaging (SCI) offers an efficient solution for high-speed video acquisition and, under exposure-time camera--scene relative motion, multi-…