Health
1353 storiesPublic health and medicine worldwide — sourced from WHO News.
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?
arXiv:2608.19799v2 Announce Type: replace Abstract: Software increasingly functions as part of the scientific instrument itself, making failures in scientific code capable of compromising not only program behavior but al…
FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
arXiv:2608.20153v2 Announce Type: replace Abstract: Large language models (LLMs) have shown growing potential for automated theoretical computer science (TCS) research, yet existing benchmarks remain far from realistic r…
Electronic Navigational Chart Change Classification
arXiv:2608.20218v2 Announce Type: replace Abstract: Electronic Navigational Charts (ENCs) are geospatial vector datasets used in maritime navigation systems that represent hydrographic and navigational information such a…
ACE-Ego-Hand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery
arXiv:2608.20308v3 Announce Type: replace Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and fr…
Neural-Primitive: An Efficient End-to-end Local Planner with Primitive-based Imitation Learning for Autonomous Flight
arXiv:2608.20948v2 Announce Type: replace Abstract: Autonomous flight in unknown cluttered environments is hindered by the computation-quality-memory trilemma of onboard trajectory generation. In this paper, we propose a…
Security Games on Series-Parallel Attack Graphs with Adaptive Attackers
arXiv:2608.21259v2 Announce Type: replace Abstract: We study security games on attack graphs, where an adaptive attacker seeks to reach a target by sequentially attempting stochastic controls along the current attack fro…
HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning
arXiv:2608.21863v2 Announce Type: replace Abstract: Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iteratively. Reinforcement Learning…
ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents
arXiv:2608.21969v3 Announce Type: replace Abstract: Humans naturally exhibit multiple forms of abstraction in reasoning and interaction, including temporal abstraction across decision timescales and strategic abstraction…
PropUQ-MAS: Propagation-Aware Uncertainty Quantification for LLM Multi-Agent Systems
arXiv:2608.22130v3 Announce Type: replace Abstract: LLM-based multi-agent systems (MAS) solve complex tasks through communication among role-specialized agents. However, inter-agent dependencies introduce reliability ris…
MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models
arXiv:2608.22236v2 Announce Type: replace Abstract: Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio …
Beyond Dense Adam States: Adaptive Log-Space Quantization for Memory-Efficient Optimizers
arXiv:2608.22322v3 Announce Type: replace Abstract: Optimizer-state quantization is commonly designed for Adam's dense, parameter-aligned first- and second-moment arrays. This abstraction breaks for memory-efficient opti…
Randomized Strategyproof Facility Location: Two Facilities and Beyond
arXiv:2608.22484v2 Announce Type: replace Abstract: We design and analyze randomized strategyproof mechanisms for multi-facility location under the utilitarian social-cost objective, the sum of the agents' distances to t…
Stress Testing Unlearning Algorithms
arXiv:2608.22527v2 Announce Type: replace Abstract: Recently, machine unlearning, the removal of specific training data influence from a model, has gained increasing attention. In large language models (LLMs), unlearning…
A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes (extended version)
arXiv:2608.23137v4 Announce Type: replace Abstract: Existing articulatory corpora based on real-time MRI and electromagnetic articulography capture tongue motion but lack traceable labels for the muscle-driven process un…
CyberFactory: Scaling Cyber Security Capabilities with Instances from the Wild
arXiv:2608.23181v3 Announce Type: replace Abstract: As large language models (LLMs) continue to advance in coding capabilities, their potential in cybersecurity has drawn increasing research attention, with closed-source…
EVEREST:Endogenous Vision-Language Reinforcement Reasoning Exploration for Urban Socio-Semantic Segmentation
arXiv:2608.24640v2 Announce Type: replace Abstract: Urban socio-semantic segmentation leverages digital and satellite imagery to provide critical spatial semantic information for downstream applications such as urban res…
RACE: Scalable Statistical Estimation of Functional Consistency in LLM Neurons
arXiv:2608.24758v2 Announce Type: replace Abstract: Discovering stable neuron behavior across entire domains remains a challenge in mechanistic interpretability. Existing methods often rely on instance-level point estima…
Latent Action as Intention Enables Efficient Future Imagination for World Action Models
arXiv:2608.24882v2 Announce Type: replace Abstract: World action models (WAMs) improve robot control by modeling how observations evolve, but generating future observations at test time incurs substantial latency. Fast-W…
Fuzzy Pattern Matching in Ordered Structures
arXiv:2608.25032v2 Announce Type: replace Abstract: The problem of pattern matching, that is, finding all occurrences of a given pattern in a string, is one of the fundamental problems in computer science that has applic…
Padamitra: Grounded Glossary Generation for Classical Sanskrit
arXiv:2608.25038v2 Announce Type: replace Abstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meani…