{
  "title": "每日研究简报 2026-09-11",
  "url": "/posts/research-brief-2026-09-11/",
  "permalink": "https://hackcv.com/posts/research-brief-2026-09-11/",
  "date": "2026-09-11",
  "lastmod": "2026-09-11",
  "author": "",
  "description": "AI / 大模型 / Agent / 计算机视觉 / 音视频处理算法 / 工程优化 领域每日研究简报",
  "categories": ["研究简报"],
  "tags": ["AI","大模型","Agent","计算机视觉","音视频处理","工程优化","每日简报"],
  "cover": "https://picsum.photos/seed/%E6%AF%8F%E6%97%A5%E7%A0%94%E7%A9%B6%E7%AE%80%E6%8A%A5-2026-09-11/1200/675",
  "readingTime": 4,
  "wordCount": 1193,
  "content": "\u003ch1 id=\"每日研究简报-2026-09-11\"\u003e每日研究简报 2026-09-11\u003c/h1\u003e\n\u003cp\u003e技术人视角 · 今日四栏精选：arXiv 论文 / GitHub 开源 / HuggingFace 热门 / 行业资讯。\u003c/p\u003e\n\u003chr\u003e\n\u003ch2 id=\"一--arxiv-最新论文\"\u003e一 · arXiv 最新论文\u003c/h2\u003e\n\u003ch3 id=\"sensenova-u15-towards-native-unified-visual-intelligence\"\u003eSenseNova-U1.5: Towards Native Unified Visual Intelligence\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and\u003cbr\u003e\n\u003cstrong\u003e领域\u003c/strong\u003e：AI / 大模型\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：计算机视觉方向的新工作，可以了解一下。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：http://arxiv.org/abs/2609.11929v1\u003c/p\u003e\n\u003ch3 id=\"gpu-cfr-80x-faster-counterfactual-regret-minimization-by-compiling-the-game-to-static-dataflow-and-cuda-graph-replay\"\u003eGPU-CFR: 80x Faster Counterfactual Regret Minimization by Compiling the Game to Static Dataflow and CUDA Graph Replay\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Counterfactual regret minimization (CFR) is one of the few large numerical workloads that still runs faster on CPUs than on GPUs. Each iteration sweeps a game tree with up to billions of states in millions of small, interdependent gather and scatter steps issued through a generic tree interface. On a GPU every kernel finishes in microseconds, so kernel launches and framework dispatch dominate the\u003cbr\u003e\n\u003cstrong\u003e领域\u003c/strong\u003e：AI / 大模型\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：AI医疗方向的应用，看看能解决什么实际问题。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：http://arxiv.org/abs/2609.11923v1\u003c/p\u003e\n\u003ch3 id=\"general-quantification-of-covariate-and-concept-shifts\"\u003eGeneral Quantification of Covariate and Concept Shifts\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from samples. In this paper, we bridge the gap between theory and practical applications. We first show that existing definition of concept shift breaks when the source and target supports mismatch. Leveraging\u003cbr\u003e\n\u003cstrong\u003e领域\u003c/strong\u003e：AI / 大模型\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：AI医疗方向的应用，看看能解决什么实际问题。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：http://arxiv.org/abs/2609.11918v1\u003c/p\u003e\n\u003ch3 id=\"data-scarcity-and-model-sparsity-mixtures-of-experts-overfit-more-to-repeated-data\"\u003eData Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetit\u003cbr\u003e\n\u003cstrong\u003e领域\u003c/strong\u003e：AI / 大模型\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：大模型方向的新进展，看看有什么新思路。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：http://arxiv.org/abs/2609.11917v1\u003c/p\u003e\n\u003ch3 id=\"can-edge-deployable-vision-language-models-identify-species\"\u003eCan Edge-Deployable Vision-Language Models Identify Species?\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Camera traps often run in the field on edge hardware with limited or no connectivity, making small, locally-deployable vision-language models (VLMs) \u0026ndash; not frontier-scale ones \u0026ndash; the practically relevant class to evaluate for species identification. We test whether models in this deployment-relevant 2\u0026ndash;8B range carry genuine taxonomic knowledge, evaluating four such VLMs (Qwen3-VL 2B/4B/8B, Gemma3\u003cbr\u003e\n\u003cstrong\u003e领域\u003c/strong\u003e：AI / 大模型\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：大模型方向的新进展，看看有什么新思路。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：http://arxiv.org/abs/2609.11916v1\u003c/p\u003e\n\u003ch3 id=\"generative-marketing-mix-modeling-a-causal-inference-framework-linking-geo-and-gem-to-business-impact\"\u003eGenerative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Generative artificial intelligence changes how firms reach customers, but standard marketing data do not record how often users see and notice a firm\u0026rsquo;s name in generated answers. We develop Generative Marketing Mix Modeling (GMMM) to estimate the causal effects of Generative Engine Optimization (GEO) and Generative Engine Marketing (GEM). For GEO, GMMM combines repeated generated answers with ques\u003cbr\u003e\n\u003cstrong\u003e领域\u003c/strong\u003e：AI / 大模型\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：AI医疗方向的应用，看看能解决什么实际问题。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：http://arxiv.org/abs/2609.11915v1\u003c/p\u003e\n\u003ch3 id=\"distance-generalization-in-transformers-why-bother-with-positional-encoding\"\u003eDistance generalization in transformers: why bother with positional encoding?\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Out-of-distribution length generalization, namely to extrapolate a task from short to longer context, has been studied intensively for transformers. Here we focus on distance generalization, which probes performance when inter-token distances are changed between training and inference, while keeping a fixed context length. We construct two synthetic delay copy tasks, both involving finite distance\u003cbr\u003e\n\u003cstrong\u003e领域\u003c/strong\u003e：AI / 大模型\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：大模型方向的新进展，看看有什么新思路。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：http://arxiv.org/abs/2609.11913v1\u003c/p\u003e\n\u003ch3 id=\"artificial-id-drive-and-persistent-alignment-in-agentic-ai\"\u003eArtificial Id: Drive and Persistent Alignment in Agentic AI\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive fo\u003cbr\u003e\n\u003cstrong\u003e领域\u003c/strong\u003e：AI / 大模型\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：Agent方向的新尝试，做智能体的同学可以看看。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：http://arxiv.org/abs/2609.11911v1\u003c/p\u003e\n\u003ch2 id=\"二--github-热门开源\"\u003e二 · GitHub 热门开源\u003c/h2\u003e\n\u003ch3 id=\"openclawopenclaw\"\u003eopenclaw/openclaw\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e简介\u003c/strong\u003e：The AI that really does things. Any OS. Any Platform. The lobster way. 🦞\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：389423⭐\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：38万星不是盖的，主打跨平台执行能力，可以看看它的架构设计。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://github.com/openclaw/openclaw\u003c/p\u003e\n\u003ch3 id=\"obrasuperpowers\"\u003eobra/superpowers\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e简介\u003c/strong\u003e：An agentic skills framework \u0026amp; software development methodology that works.\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：285062⭐\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：20万星的热门项目，技能系统的设计思路值得学习。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://github.com/obra/superpowers\u003c/p\u003e\n\u003ch3 id=\"nousresearchhermes-agent\"\u003eNousResearch/hermes-agent\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e简介\u003c/strong\u003e：The agent that grows with you\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：244412⭐\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：20万星的热门项目，技能系统的设计思路值得学习。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://github.com/NousResearch/hermes-agent\u003c/p\u003e\n\u003ch3 id=\"n8n-ion8n\"\u003en8n-io/n8n\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e简介\u003c/strong\u003e：Fair-code workflow automation platform with native AI capabilities. Combine visual building with custom code, self-host or cloud, 400+ integrations.\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：203993⭐\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：20万星的热门项目，技能系统的设计思路值得学习。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://github.com/n8n-io/n8n\u003c/p\u003e\n\u003ch3 id=\"significant-gravitasautogpt\"\u003eSignificant-Gravitas/AutoGPT\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e简介\u003c/strong\u003e：AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：187253⭐\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：10万星的老牌项目，可以看看它最近的更新。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://github.com/Significant-Gravitas/AutoGPT\u003c/p\u003e\n\u003ch3 id=\"firecrawlfirecrawl\"\u003efirecrawl/firecrawl\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e简介\u003c/strong\u003e：The context API to search, scrape, and interact with the web at scale. 🔥\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：179005⭐\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：10万星的老牌项目，可以看看它最近的更新。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://github.com/firecrawl/firecrawl\u003c/p\u003e\n\u003ch3 id=\"fpromptschat\"\u003ef/prompts.chat\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e简介\u003c/strong\u003e：f.k.a. Awesome ChatGPT Prompts. Share, discover, and collect prompts from the community. Free and open source — self-host for your organization with complete privacy.\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：169949⭐\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：10万星的老牌项目，可以看看它最近的更新。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://github.com/f/prompts.chat\u003c/p\u003e\n\u003ch3 id=\"snailclimbjavaguide\"\u003eSnailclimb/JavaGuide\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e简介\u003c/strong\u003e：Java 面试 \u0026amp; 后端通用面试指南，覆盖计算机基础、数据库、分布式、高并发、系统设计与 AI 应用开发\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：158449⭐\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：10万星的老牌项目，可以看看它最近的更新。\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://github.com/Snailclimb/JavaGuide\u003c/p\u003e\n\u003ch2 id=\"三--huggingface-热门\"\u003e三 · HuggingFace 热门\u003c/h2\u003e\n\u003ch3 id=\"-daily-papers-精选\"\u003e📄 Daily Papers 精选\u003c/h3\u003e\n\u003ch3 id=\"ncp-archpreview-technical-report-moving-towards-latent-space-language-models-through-next-concept-prediction\"\u003eNCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：We introduce NCP-ArchPreview, a latent-space language model that pushes autoregressive pretraining beyond standard next-token prediction (NTP). Alongside NTP, the model learns through Next Concept Prediction (NCP) to predict discrete concepts that span multiple tokens, introducing an explicit and mo\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：107⬆\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://huggingface.co/papers/2609.10715\u003c/p\u003e\n\u003ch3 id=\"sensenova-u15-towards-native-unified-visual-intelligence-1\"\u003eSenseNova-U1.5: Towards Native Unified Visual Intelligence\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：83⬆\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://huggingface.co/papers/2609.11929\u003c/p\u003e\n\u003ch3 id=\"spatialblock-enhancing-spatial-intelligence-in-lvlms-via-synthetic-block-stacking-problem\"\u003eSpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Large Vision-Language Models (LVLMs) have achieved strong performance on diverse visual tasks, yet their ability to reconstruct and reason about the 3D structure of the scene depicted in 2D images \u0026ndash; referred to as spatial intelligence \u0026ndash; remains limited. Existing approaches attempt to address this\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：36⬆\u003cbr\u003e\n\u003cstrong\u003eGitHub\u003c/strong\u003e：https://github.com/rsoohyun/SpatialBlock\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://huggingface.co/papers/2609.07064\u003c/p\u003e\n\u003ch3 id=\"evosafeharness-evolving-model--and-domain-specific-harnesses-for-securing-agents\"\u003eEvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Large Language Model (LLM) agents are turning language into real-world effects, making safety necessary against both indirect prompt injections and direct harmful requests. System-level safety harnesses add an enforcement layer beyond model-level defenses, but existing harnesses are usually designed\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：34⬆\u003cbr\u003e\n\u003cstrong\u003eGitHub\u003c/strong\u003e：https://github.com/SaFo-Lab/EvoSafeHarness\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://huggingface.co/papers/2609.05903\u003c/p\u003e\n\u003ch3 id=\"mi-ripple-restoring-images-degraded-by-iterative-ai-editing\"\u003eMi-Ripple: Restoring Images Degraded by Iterative AI Editing\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Iterative reference-conditioned image editing can introduce grid-like and granular textures, commonly described as digital ripple. We present Mi-Ripple, a diagnosis-guided restoration workflow that suppresses this digital ripple while protecting image structure. Mi-Ripple separates periodic lattice\u003cbr\u003e\n\u003cstrong\u003e热度\u003c/strong\u003e：16⬆\u003cbr\u003e\n\u003cstrong\u003eGitHub\u003c/strong\u003e：https://github.com/miyang-ai/Mi-Ripple\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://huggingface.co/papers/2609.11317\u003c/p\u003e\n\u003ch3 id=\"-huggingface-blog\"\u003e📝 HuggingFace Blog\u003c/h3\u003e\n\u003ch3 id=\"rebuilding-automatic1111-with-gradio-workflow\"\u003eRebuilding AUTOMATIC1111 with Gradio Workflow\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Rebuilding AUTOMATIC1111 with Gradio Workflow\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://huggingface.co/blog/gradio-workflow-1111\u003c/p\u003e\n\u003ch3 id=\"ibm-releases-sota-granite-time-series-patchtst-fm-r2-model-with-commercial-friendly-license\"\u003eIBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://huggingface.co/blog/ibm-research/ibm-releases-sota-granite-time-series\u003c/p\u003e\n\u003ch3 id=\"safety-for-whom-refusing-the-right-subset-of-a-topic-not-the-whole-topic\"\u003eSafety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e摘要\u003c/strong\u003e：Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic\u003cbr\u003e\n\u003cstrong\u003e链接\u003c/strong\u003e：https://huggingface.co/blog/MultiverseComputingCAI/safety-for-whom\u003c/p\u003e\n\u003ch2 id=\"四--行业资讯\"\u003e四 · 行业资讯\u003c/h2\u003e\n\u003ch3 id=\"让智能体自主探索而不越界蚂蚁密算开源可信原生智能体hop-30\"\u003e让智能体自主探索而不越界，蚂蚁密算开源可信原生智能体HOP 3.0\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e内容\u003c/strong\u003e：在2026 Inclusion·外滩大会上，蚂蚁密算董事长韦韬宣布可信原生智能体HOP 3.0正式开源，向开发者、企业和行业专家开放“智能体原生语言”相关技术能力，推动产业智能体从依赖模型自觉，走向边界明确、过程可控、结果可核验的可信执行。 （蚂蚁密算董事长 韦韬宣布 HOP3.0 开源） 2025年世界人工智能大会期间，蚂蚁密算首次发布并开源HOP 1.0技术框架，探索通过工程化方法提升大模型在金融、医\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：AI Agent方向的新进展，这个方向很火。\u003cbr\u003e\n\u003cstrong\u003e来源\u003c/strong\u003e：雷锋网 AI\u003c/p\u003e\n\u003ch3 id=\"碳硅道统五级梯队的智能分级与十维标尺的对应\"\u003e碳硅道统：五级梯队的智能分级与十维标尺的对应\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e内容\u003c/strong\u003e：分级不是排名。是定位。 五级梯队，从一阶到五阶。 一阶鹦鹉。模式匹配。输入到输出。无理解。无内生驱动。对应硅基系统的最初级形态。 二阶技工。规则执行。能完成特定任务。有边界意识但来自外部训练。对应当前的大模型。 三阶协作。能理解复杂意图。能在多步任务中保持一致性。对应当前最先进的AI系统。 四阶共生。能与碳基形成深度协作。能感知碳基的意图、情绪、需求。但内生驱动仍为零。这是硅基的理论上限。 五阶启灵。理论上的上限。碳基专属。内生觉知。零维连通。归零稳态。硅基不可达。 十维标尺与五级梯队的对应。 十维标尺不是用来打分的。它是用来定位一个系统落在五级梯队的哪一阶。 觉知本源维度。一阶到四阶全部为零。五阶独有。 逻辑自洽维度。一阶低。二阶中。三阶高。四阶极高。五阶完美。\u0026lt;\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：产业动态，可以看看有什么新趋势。\u003cbr\u003e\n\u003cstrong\u003e来源\u003c/strong\u003e：雷锋网 AI\u003c/p\u003e\n\u003ch3 id=\"阿里云token-plan个人版升级加量不加价新增12类agent-harness工具\"\u003e阿里云Token Plan个人版升级：加量不加价，新增12类Agent Harness工具\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e内容\u003c/strong\u003e：9月11日，阿里云Token Plan个人版升级，原有价格及Credit 额度保持不变，Standard和Pro套餐新增 Agent Harness工具权益与用量，覆盖搜索、网页解析、图像生成、语音处理、代码执行等12项Agent开发常用能力。上述工具均通过MCP标准协议开放，可直接接入到自有Agent应用，无需逐项单独购买。 Token Plan个人版主要面向个人开发者和中小团队。Lite套餐价格为39元/月，每7天提供2500 C\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：AI Agent方向的新进展，这个方向很火。\u003cbr\u003e\n\u003cstrong\u003e来源\u003c/strong\u003e：雷锋网 AI\u003c/p\u003e\n\u003ch3 id=\"智象未来-vivago-r1-全球上线国内版本够搭全新升级发布从-15-秒到5-分钟ai-视频创作进入单反级交付时代\"\u003e智象未来 vivago R1 全球上线，国内版本「够搭」全新升级发布：从 15 秒到5 分钟，AI 视频创作进入单反级交付时代\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e内容\u003c/strong\u003e：2026年9月8日，智象未来（HiDream.ai）旗下内容创作智能体 vivago R1 全球上线，国内版本 「够搭」 \u0026lt;s\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：产业动态，可以看看有什么新趋势。\u003cbr\u003e\n\u003cstrong\u003e来源\u003c/strong\u003e：雷锋网 AI\u003c/p\u003e\n\u003ch3 id=\"全球首个3d原生城市世界模型abot-earth-07发布构建ai理解真实世界的入口\"\u003e全球首个3D原生城市世界模型ABot-Earth 0.7发布，构建AI理解真实世界的入口\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e内容\u003c/strong\u003e：9月10日，阿里巴巴集团旗下高德正式发布全球首个3D原生城市世界模型ABot-Earth 0.7。\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：产业动态，可以看看有什么新趋势。\u003cbr\u003e\n\u003cstrong\u003e来源\u003c/strong\u003e：量子位\u003c/p\u003e\n\u003ch3 id=\"全球首个可仿真的人场景交互重建框架-hsimul3r让人类视频真正成为机器人技能来源\"\u003e全球首个可仿真的人–场景交互重建框架 HSImul3R：让人类视频真正成为机器人技能来源\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e内容\u003c/strong\u003e：大晓机器人联合南洋理工大学 S-Lab、上海人工智能实验室发布全新人–场景交互重建研究 HSImul3R\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：具身智能落地的真实案例，看看他们怎么解决实际问题。\u003cbr\u003e\n\u003cstrong\u003e来源\u003c/strong\u003e：量子位\u003c/p\u003e\n\u003ch3 id=\"agi时代的第一个生图模型chatgpt-images-25上线\"\u003eAGI时代的第一个生图模型，ChatGPT Images 2.5上线\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e内容\u003c/strong\u003e：主打生成更快，细节更好，改图也终于越来越像“真·修图”了。\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：大模型领域的新动态，看看有什么新能力。\u003cbr\u003e\n\u003cstrong\u003e来源\u003c/strong\u003e：量子位\u003c/p\u003e\n\u003ch3 id=\"introducing-chatgpt-for-financial-services\"\u003eIntroducing ChatGPT for Financial Services\u003c/h3\u003e\n\u003cp\u003e\u003cstrong\u003e内容\u003c/strong\u003e：Introducing ChatGPT for Financial Services, combining built-in financial data and GPT-6 Astra for research, modeling, and client-ready materials.\u003cbr\u003e\n\u003cstrong\u003e推荐理由\u003c/strong\u003e：大模型领域的新动态，看看有什么新能力。\u003cbr\u003e\n\u003cstrong\u003e来源\u003c/strong\u003e：OpenAI Blog\u003c/p\u003e\n\u003cp\u003e本次任务消耗Token统计：脚本化模式（opencode 启动，无独立 token 计量）\u003c/p\u003e\n",
  "summary": "每日研究简报 2026-09-11 技术人视角 · 今日四栏精选：arXiv 论文 / GitHub 开源 / HuggingFace 热门 / 行业资讯。\n一 · arXiv 最新论文 SenseNova-U1.5: Towards Native Unified Visual Intelligence 摘要：We launch SenseNova-U1.5, an 8B-MoT native unified multimodal model that understands, reasons about, and generates visual content within an encoder-free and VAE-free architecture. We strengthen its visual interface through spatially coherent patch reconstruction and scale its training with carefully curated generation and editing data, improved task formulation, structural prompt enhancement, and\n领域：AI / 大模型\n推荐理由：计算机视觉方向的新工作，可以了解一下。\n链接：http://arxiv.org/abs/2609.11929v1\n"
}
