📑 Table of Contents

Daily Research Brief 2026-07-07

📊 Token usage: estimated from retrieval and writing scale.

Covers the latest AI research, open source and industry moves, updated daily.


Editor’s Note

WAIC 2026 is set for July 17-20 in Shanghai with major releases expected; Meituan open-sourced LongCat-2.0 (1.6T parameters), closing the loop on trillion-scale training with domestic compute. On the paper side, HAS-Bench, AgentGym2 and CausalGame keep pushing agent evaluation toward real-world, human-in-the-loop settings.

1. Latest arXiv Papers

  1. HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participationhttps://arxiv.org/abs/2607.04329

  2. Toward Trustworthy Large Language Model Agents in Healthcarehttps://arxiv.org/abs/2607.05055

  3. Latent Programming Horizons in Coding Agentshttps://arxiv.org/abs/2607.05188

  4. LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RLhttps://arxiv.org/abs/2607.04412

  5. Rethinking On-Policy Self-Distillation for Thinking Modelshttps://arxiv.org/abs/2607.05184

  6. CausalGame: Benchmarking Causal Thinking of LLM Agents in Gameshttps://arxiv.org/abs/2607.04293

  7. AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environmentshttps://arxiv.org/abs/2607.05174

  8. Program-as-Weights: A Programming Paradigm for Fuzzy Functionshttps://arxiv.org/abs/2607.02512

Join the discussion

Comments powered by GitHub Discussions, stored in the hackcv/blog repo; sign in with a GitHub account to join. Markdown and emoji supported.

Comments powered by GitHub Discussions, stored in the hackcv/blog repo; sign in with a GitHub account to join. Markdown and emoji supported.