{
  "title": "Daily Research Brief 2026-07-07",
  "url": "/en/posts/research-brief-2026-07-07/",
  "permalink": "https://hackcv.com/en/posts/research-brief-2026-07-07/",
  "date": "2026-07-07",
  "lastmod": "2026-07-07",
  "author": "",
  "description": "Daily research brief — AI / LLM / Agent / Computer Vision / Audio-Video / Engineering",
  "categories": ["Research Brief"],
  "tags": ["AI","LLM","Agent","Computer Vision","Audio-Video","Engineering","Daily Brief"],
  "cover": "https://picsum.photos/seed/daily-research-brief-2026-07-07/1200/675",
  "readingTime": 1,
  "wordCount": 152,
  "content": "\u003ch1 id=\"daily-research-brief-2026-07-07\"\u003eDaily Research Brief 2026-07-07\u003c/h1\u003e\n\u003cp\u003e📊 Token usage: estimated from retrieval and writing scale.\u003c/p\u003e\n\u003cp\u003eCovers the latest AI research, open source and industry moves, updated daily.\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"editors-note\"\u003eEditor\u0026rsquo;s Note\u003c/h2\u003e\n\u003cp\u003eWAIC 2026 is set for July 17-20 in Shanghai with major releases expected; Meituan open-sourced LongCat-2.0 (1.6T parameters), closing the loop on trillion-scale training with domestic compute. On the paper side, HAS-Bench, AgentGym2 and CausalGame keep pushing agent evaluation toward real-world, human-in-the-loop settings.\u003c/p\u003e\n\u003ch2 id=\"1-latest-arxiv-papers\"\u003e1. Latest arXiv Papers\u003c/h2\u003e\n\u003col\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eHAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2607.04329\"\u003ehttps://arxiv.org/abs/2607.04329\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eToward Trustworthy Large Language Model Agents in Healthcare\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2607.05055\"\u003ehttps://arxiv.org/abs/2607.05055\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eLatent Programming Horizons in Coding Agents\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2607.05188\"\u003ehttps://arxiv.org/abs/2607.05188\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eLLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2607.04412\"\u003ehttps://arxiv.org/abs/2607.04412\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eRethinking On-Policy Self-Distillation for Thinking Models\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2607.05184\"\u003ehttps://arxiv.org/abs/2607.05184\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eCausalGame: Benchmarking Causal Thinking of LLM Agents in Games\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2607.04293\"\u003ehttps://arxiv.org/abs/2607.04293\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eAgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2607.05174\"\u003ehttps://arxiv.org/abs/2607.05174\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003cli\u003e\n\u003cp\u003e\u003cstrong\u003eProgram-as-Weights: A Programming Paradigm for Fuzzy Functions\u003c/strong\u003e — \u003ca href=\"https://arxiv.org/abs/2607.02512\"\u003ehttps://arxiv.org/abs/2607.02512\u003c/a\u003e\u003c/p\u003e\n\u003c/li\u003e\n\u003c/ol\u003e\n",
  "summary": "Daily Research Brief 2026-07-07 📊 Token usage: estimated from retrieval and writing scale.\nCovers the latest AI research, open source and industry moves, updated daily.\nEditor\u0026rsquo;s Note WAIC 2026 is set for July 17-20 in Shanghai with major releases expected; Meituan open-sourced LongCat-2.0 (1.6T parameters), closing the loop on trillion-scale training with domestic compute. On the paper side, HAS-Bench, AgentGym2 and CausalGame keep pushing agent evaluation toward real-world, human-in-the-loop settings.\n"
}
