Qortora · Search · Indexed page

www.alphaxiv.orgFetched 2026-08-17T10:59:40Z

Explore | alphaXiv

Discuss, discover, and read arXiv papers. Explore trending papers, see recent activity and discussions, and follow authors of arXiv papers on alphaXiv.

Open original source · Full cached text

Explore | alphaXiv alphaXiv Explore Researchers Sign In MCP Server Autoresearch Browser Extension Light themeDark theme BlogSend Feedback? Ask questions across all of research alphaXiv connects papers, researchers, and organizations, grounding the answer in the underlying work. What's worth reading?Discover relevant new papers, curated by our research communityWho's working on it?Surface the researchers and institutions behind a topicGrounded literature reviewGet evidence-based answers through in-line citations to papers Alt + Enter to search Sign up Personalize your feed GLM-5.3: Frontier Coding with Emergent Cyber Capabilities 14 Aug 2026Z.ai GLM-5.3 is built on the same base model as GLM-5.2, with every reported gain coming from scaled post-training on long-horizon task environments. It posts open-source SOTA on Terminal Bench 3.0 (28.3 vs 4.6) and Agents' Last Exam (28.5), and a 50% improvement on Z.ai's in-house Code Bench while spending fewer output tokens. Cyber capability grew fastest of all: 84.5% on CyberGym is the best result on that benchmark, and exploitation scores more than doubled. Weights are slated for release two weeks after launch, once safety hardening completes. 192 Bookmark Training AI Scientists to Replicate Research 13 Aug 2026Damon Falck Samer SabriAnja Surina Researchers at Inherent developed Faraday, an AI Scientist agent designed to replicate research papers by reproducing experimental figures using a "Coding Agent as a Tool" (CAT) paradigm. Trained on the Replica task space with a novel rubric-based reward system, Faraday demonstrated superior scientific rigor and replication accuracy compared to leading frontier models, including Claude Opus 4.8 and GPT-5.5. 50 Bookmark Autoresearch View PDF793 Latent On-Policy Self-Distillation 13 Aug 2026Guibin ZhangJiayang LyuRan Sun Latent On-Policy Self-Distillation (LOPD) introduces a framework where the privileged context for self-distillation is a learnable latent representation, moving beyond human-engineered artifacts. This approach consistently enhanced agent performance on tool-use and code generation benchmarks, yielding superior aggregate results and improved sample efficiency across multiple LLM backbones. 40 Bookmark Autoresearch 10 View PDF549 Researchers to follow View all Yann LeCun Executive Chairman @ AMI - Advanced Machine Intelligence, Jacob T. Schwartz Professor, CS @ New York University Follow Ilya Sutskever CEO and Co-Founder @ Safe Superintelligence Inc, Previously Co-Founder and Chief Scientist @ OpenAI Follow Saining Xie Co-Founder and CSO @ AMI Labs, Assistant Professor, CS @ New York University Follow Sergey Levine Co-Founder @ Physical Intelligence, Associate Professor, EECS @ UC Berkeley Follow Yilun Du Assistant Professor, CS @ Harvard University, Institute Investigator @ Kempner Institute Follow Andrew Ng Managing Partner @ AI Aspire, Managing General Partner @ AI Fund, Executive Chairman @ LandingAI, Founder @ DeepLearning.AI, Adjunct Professor, CS @ Stanford University, Chairman and Co-Founder @ Coursera Follow Kaiming He Distinguished Scientist @ Google DeepMind, Associate Professor, EECS @ MIT Follow Yoshua Bengio President and Scientific Director @ LawZero, Founder and Scientific Advisor @ Mila - Quebec Artificial Intelligence Institute, Canada CIFAR AI Chair @ CIFAR, Full Professor, CS @ Université de Montréal Follow Small-Scale Experiments: Are We There Yet? 12 Aug 2026Nicholas LourieKyunghyun ChoKaren Ullrich This research from FAIR at MSL Meta and New York University establishes that reliable scaling laws are observable even in very small foundation models, provided hyperparameters are rigorously tuned. The study presents a methodology that enables cost-effective, model-centric research by revealing how hyperparameter sensitivity at smaller scales can obscure these laws. 68 Bookmark Autoresearch View PDF1,247 Flux-Form Spatiotemporal Neural Operators for Coarse-Grained Dynamics of Multiscale PDEs 16 Aug 2026Junfeng Chen A new class of Flux-Form Spatiotemporal Neural Operators enables stable and accurate coarse-grained predictions for multiscale PDE systems by explicitly embedding local conservation laws and causal history dependence. The operators consistently reproduce long-horizon dynamics and time-averaged statistics, surpassing both traditional physics-based models and purely data-driven approaches across Burgers', Kuramoto-Sivashinsky, and Navier-Stokes equations. 2 Bookmark Autoresearch View PDF34 AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design 13 Aug 2026Yaxin LuoHaobin JiangJialv Zou Meituan and MBZUAI researchers developed AutoDesign, a meta-harness optimization framework enabling design systems to recursively improve their operational components based on human-aligned evaluations. Applied to academic paper-to-poster generation, it achieved a PosterBench Score of 78.32, outperforming other systems by over 7 points, and autonomously generated posters in 40 minutes at under $3 each. 39 Bookmark Autoresearch 44 View PDF481 Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development 13 Aug 2026Yiwei LiWanli YangHexiang Tan A systematic evaluation framework was developed for autonomous AI R&D agents, moving beyond single final scores to diagnose performance using process-level metrics, controlled comparisons for experience reuse, and harness impact analysis. The study revealed that current frontier models reliably optimize artifacts but exhibit rare genuine innovation (1.2% novel solutions) and inconsistent performance heavily influenced by bottlenecks in solution framing and feedback control, as well as the effectiveness of experience reuse. 18 Bookmark Autoresearch View PDF208 DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation 13 Aug 2026DreamX TeamRui ChenXiangxiang Chu DreamX-Phi 1.0 introduces an action-conditioned video world model for robotic manipulation, incorporating geometry-aware SE(3) action representation and comprehensive physical-consistency supervision. This model achieved first place on the WorldArena 2.0 Track 1 leaderboard for video prediction with an EWMScore-P of 60.65 and tied for second on Track 2 for policy training, demonstrating a 67.19% success rate on the "Adjust Bottle" task. 37 Bookmark Autoresearch 45 View PDF384 OmniScientist: An Omni-Modal Omni-Discipline AI Scientist 13 Aug 2026Bobo LiHao FeiTianjie Ju OmniScientist presents an end-to-end AI scientist capable of conducting multidisciplinary research directly from raw, heterogeneous scientific evidence, integrating perception throughout the research lifecycle. It successfully completes full research workflows across diverse scientific domains and modalities, demonstrating improved research quality and multimodal grounding compared to systems relying on pre-processed data. 18 Bookmark Autoresearch View PDF181 V-RAE: Rethinking Video Latent Spaces for Generation 13 Aug 2026Minghui GuoShengqiong Wu Hao Fei V-RAE constructs generative latent spaces for video by integrating frozen Vision Foundation Models (VFMs) with a learnable temporal pooling module and a spatiotemporal decoder. This method improves video generation quality, reduces diffusion model training time by up to 6x, and maintains strong semantic information in the latent representations. 27 Bookmark Autoresearch View PDF261 Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding 11 Aug 2026Kushal Chakrabarti Agent instruction files like CLAUDE.md exhibit unbounded growth due to "catastrophic remembering," a process where the original rationale for instructions is lost, making safe deletion difficult. Implementing prompt comments that record outcome-grounded reasoning effectively halts this growth, reducing excess instruction size from 211.3% to 1.4% and improving agent instruction-following correctness by 11.6 percentage points. 111 Bookmark Autoresearch View PDF1,624 Researchers to follow View all Stefano Ermon CEO & Co-Founder @ Inception, Associate Professor, CS @ Stanford University Follow Dario Amodei CEO and Co-Founder @ Anthropic, Previously Vice President of Research @ OpenAI Follow Yejin Choi The Dieter Schwarz Foundation Professor, CS & Senior Fellow, HAI @ Stanford University, Distinguished Scientist, Language and Cognition Research @ NVIDIA Follow Dawn Song Vice President of AI Research @ Meta, Professor, CS @ UC Berkeley Follow Jingren Zhou Chief Scientist @ Alibaba Group, Previously Researcher @ Microsoft Research Follow Noam Shazeer Previously VP Engineering @ Google Follow Junyang Lin Independent Researcher @ Unaffiliated, Previously Tech Lead @ Alibaba Group Follow Oriol Vinyals Co-Founder @ Discovery Loop, Previously VP of Research @ Google DeepMind Follow SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation 14 Aug 2026Jinsheng QuanJianhua LiSiyi Xie The SPARGen model from Zhejiang University and SenseTime Research unifies 3D reconstruction, dense correspondence, and spatial reasoning by formulating these as instruction-conditioned generation tasks within a single multimodal generative framework. It achieves competitive performance on visual geometry and optical flow benchmarks, and leads open-source models in spatial reasoning tasks, demonstrating improved understanding through integrated spatial supervisions. 1 Bookmark Autoresearch View PDF37 Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning 14 Aug 2026Kai ChenJifeng DingNing Ding The Intern-S2-Mobius research presents Mobius-v0, an architectural paradigm that separates knowledge storage from reasoning operations into a shared global memory. This approach allows the model to achieve comparable performance to Transformer baselines with 62.6% less training data and accelerate inference throughput by up to 4 times, primarily by producing significantly shorter and more concise reasoning outputs. 2 Bookmark Autoresearch View PDF32 PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives 14 Aug 2026Kaixin DingXi ChenMinghong Cai PlayWorld introduces a benchmark for evaluating interactive video world models by employing multi-modal Agent Players to pursue long-horizon objectives and adaptively adjust action sequences. The evaluation, performed on nine models, indicates that current systems face significant challenges in maintaining sustained world evolution and global spatial consistency across complex interactions. 15 Bookmark Autoresearch View PDF107 StreamTTT: Reconciling Real-Time Perception and Long-Term Memory in Streaming VLMs 13 Aug 2026Joya ChenZeyun ZhongMike Zheng Shou STREAMTTT introduces a streaming Video-Language Model that overcomes the perception-memory trade-off in real-time video understanding through a dual-memory architecture and a dedicated real-time QA corpus. Its 4B parameter model achieved a 68.59 two-track average on OVO-Bench, outperforming HERMES-7B by 9.39 points and improving real-time perception by 1.4 points and backward tracing by 3.7 points compared to a matched-scale baseline. 22 Bookmark Autoresearch View PDF159 Marionette: Predicting World States, Rendering Geometry, Painting Appearance 14 Aug 2026Zian MengZhen LiChuanhao Li Interactive game world models typically autoregress visual observations directly in pixel or latent space, forcing structured properties such as pose, geometry, and occlusion to be implicitly maintained by the same generative sequence. Over long horizons, errors in these latent world properties accumulate, making consistency and controllability fragile. We explicitly model the evolving world state, delegate exact geometric computation to a fixed, zero-parameter renderer, and leave the neural model to synthesize appearance. We instantiate this idea as Marionette, a world model for interactive games with articulated characters. First, a two-stage autoregressive dynamics model predicts an explicit and interp…