📑 Table of Contents
Daily Research Brief 2026-07-07
📊 Token usage: estimated from retrieval and writing scale.
Covers the latest AI research, open source and industry moves, updated daily.
Editor’s Note
WAIC 2026 is set for July 17-20 in Shanghai with major releases expected; Meituan open-sourced LongCat-2.0 (1.6T parameters), closing the loop on trillion-scale training with domestic compute. On the paper side, HAS-Bench, AgentGym2 and CausalGame keep pushing agent evaluation toward real-world, human-in-the-loop settings.
1. Latest arXiv Papers
-
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation — https://arxiv.org/abs/2607.04329
-
Toward Trustworthy Large Language Model Agents in Healthcare — https://arxiv.org/abs/2607.05055
-
Latent Programming Horizons in Coding Agents — https://arxiv.org/abs/2607.05188
-
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL — https://arxiv.org/abs/2607.04412
-
Rethinking On-Policy Self-Distillation for Thinking Models — https://arxiv.org/abs/2607.05184
-
CausalGame: Benchmarking Causal Thinking of LLM Agents in Games — https://arxiv.org/abs/2607.04293
-
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments — https://arxiv.org/abs/2607.05174
-
Program-as-Weights: A Programming Paradigm for Fuzzy Functions — https://arxiv.org/abs/2607.02512
Join the discussion
Comments powered by GitHub Discussions, stored in the hackcv/blog repo; sign in with a GitHub account to join. Markdown and emoji supported.Comments powered by GitHub Discussions, stored in the hackcv/blog repo; sign in with a GitHub account to join. Markdown and emoji supported.