We need answer in Chinese, markdown. Need follow user requirements. Need produce knowledge card draft. Need include title first line with # 知识卡片:中文标题. Need sections. Need not wrap code block. Need not fabricate; unconfirmed = 待核实. Need preserve English title, keywords, original source. Need maybe include "一句话结论", "事件概述或研究问题", "方法/产品要点", "主要结果或产业意义", "为什么重要", "局限与不确定性", "可用于图书/PPT/简报的角度", "原始材料". Also mention "与既有脉络的关系" maybe in why important. Need avoid duplicate facts from existing cards. Need not copy phrasing. Need include incremental info. Need based on source material. Source is arxiv abstract. Need craft.
Need think carefully.
We have source with arxiv ID 2608.20318v1. Title AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement. Authors etc. Published 2026-08-20. Need note "待核实" for things not in abstract? Need perhaps "作者列表、发布日期" from source. Need preserve English title and keywords. Need choose Chinese title maybe "知识卡片:AI4AI-Bench:评测大语言模型智能体在递归自我改进中的算法设计能力". Need maybe call "AI4AI-Bench".
Need include "英文关键词": perhaps "recursive self-improvement; algorithmic design; LLM agents; benchmark; training algorithms" (from abstract). Need "原始材料": URL.
Need perhaps "一句话结论": AI4AI-Bench = 10个冻结研究仓库/10类训练算法族;智能体4小时改写训练算法,重新训练最多12小时并用固定评估器打分。6系统29配置平均分0.166,最优0.250;说明LLM智能体目前只能逼近"从已有算法到最优"差距不到1/5。 Most submissions never change learning algorithm at all; those that do perform better; more reasoning effort increases willingness to modify learning algorithm and mean score. Need maybe "recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems." Need "process is training algorithm: better objective or update rule improves compute-capability exchange rate for every subsequent run." Need "No benchmark isolates that ability; existing suites won by data collection or hyperparameter tuning; none tells change to run execution apart from change to model learns." Need.
Need avoid duplicating existing cards? They are about PAST-Bench and Frontis etc. This card specifically AI4AI-Bench. Need mention "与既有脉络的关系" maybe "与 PAST-Bench 不同,AI4AI-Bench 不是测个人智能体从经验中改进,而是测LLM智能体能否设计/改写训练算法本身;与有界自我精炼分类法相比,它为RSI提供可重复的量化基准." Need only if relevant. Need not overdo. The instruction says "参考已有相关卡片,避免重复陈述同一事实;如果本条是对它们的延续或更新,请在“与既有脉络的关系”或“为什么重要”中简要说明增量信息." We can include "与既有脉络的关系" section maybe between why important and limitation. Need not copy wording. Need mention "增量信息:AI4AI-Bench 直接测量算法设计能力,而不只是自我精炼或经验保留;并公开所有打分提交。"
Need ensure "中文正文" and "保留英文标题、英文关键词和原始来源" means include English title maybe in "原始材料" or opening. Need "英文关键词" as a section. Need "不要编造事实;无法从材料确认的部分写待核实" - Need decide if "published 2026-08-20" is confirmed. Yes. "B300" is ambiguous? It says B300, likely NVIDIA? Need not say NVIDIA if not in source. Mention "单块B300(未说明具体厂商/型号细节,待核实)"? Need maybe "one B300" in source; if we mention "B300 GPU" could be okay? B300 could be some compute platform. To avoid fabrication, say "单块 B300" as in source. Don't say GPU unless source says? It says "one B300" not "GPU"; but in context likely compute accelerator. Need "待核实" maybe "B300 的具体产品类型/配置未在摘要中说明(待核实)" to be safe.
Need include "方法/产品要点" section with bullet points. Need "产品要点" maybe "AI4AI-Bench 是一个基准测试套件(benchmark suite),包含10个冻结的研究仓库、10个训练算法族;每个任务:agent 4小时内重写训练算法;随后从零重跑最多12小时;由agent不可见的固定评估器按同一流程评分,与原仓库算法比较." Need "由于10个任务的指标不可公度,每个任务映射到同一尺度:0=无信息模型,0.1=仓库自带算法,1.0=任务最优." Need "29 configurations of 6 systems on all 10 tasks" maybe "6个系统×10任务共29种配置" Wait "Across 29 configurations of 6 systems on all 10 tasks" Means 29 configurations of 6 systems? Need parse carefully. "Across 29 configurations of 6 systems on all 10 tasks the mean score is 0.166" Could mean evaluated 29 configurations? Need "6 systems" might be base agents; 29 configurations could be combinations. Need not overclarify. Need say "在全部10个任务上,6个系统的29种配置平均得分0.166" from source. Need "best system reaches 0.250" "even strongest closes under a fifth of the distance between existing algorithm and optimum." Actually "under a fifth of the distance between the algorithm that was already there and the optimum" Wait if 0.1 is repository algorithm, 1.0 optimum, max score 0.250 means (0.250-0.1)/(1.0-0.1)=0.1667? Under a fifth of distance yes. Need phrase: "即使最强系统也只完成从仓库基线(0.1)到最优(1.0)距离的不到1/5." Let's calculate: 0.25 - 0.1 = 0.15; distance to optimum = 0.9; 0.15/0.9 = 16.7%, less than 20%. Good.
Need "submissions show where distance went: most never change how model learns at all; minority that do average 0.226 vs 0.126 for rest." Need mention "The minority that do [change how model learns] average 0.226 against 0.126 for the rest." Need "More reasoning effort mostly buys willingness to go there, taking that minority from 8% of submissions to 64% and mean score from 0.094 to 0.196." Need "higher reasoning effort" maybe "更多推理努力" - source says "More reasoning effort mostly buys the willingness to go there" - good.
Need "主要结果或产业意义" section. Need "主要结果" include scores; "产业意义" maybe "为RSI提供可重复的评测协议;当前LLM智能体在算法设计上表现有限;差距主要不在于推理能力,而在于是否会去修改学习算法本身." Need ensure no fabricated "not reasoning ability" - Source says "More reasoning effort mostly buys willingness to go there" not exactly "not reasoning ability"; we can say "摘要将更多推理努力的效果主要归因于‘愿意去修改’,而不是直接保证更好的算法设计" maybe.
Need "局限与不确定性" include "摘要未提供任务具体复杂度、基线模型细节、评估器指标设计、计算成本等;需读全文核实." Need "B300 exact details, 6 system names, 29 configurations specifics, evaluation metric scale anchor '0.1' from repository algorithm" etc. Need "待核实" explicitly. Need maybe "结果基于2026-08-20 arXiv v1,尚未经过同行评议(待核实)" Wait source doesn't say peer review status. arXiv isn't peer reviewed. But "peer review status" not confirmed; can say "是否为同行评议版本:待核实(arXiv预印本)" Is that okay? We can say "来自arXiv预印本;同行评议状态未在摘要中说明(待核实)." That's okay.
Need "为什么重要" include "RSI feasibility hinges on whether agent can design training algorithms. Existing benchmarks won by data collection/hyperparameter tuning; none isolates change in how run executes from change in how model learns. AI4AI-Bench attempts to isolate this ability." Need "open-source release: task suite, evaluators, every scored submission; measurement repeatable as systems change." Need "与既有脉络" maybe "与PAST-Bench存在本质不同..." Need not duplicate.
Need "可用于图书/PPT/简报的角度" maybe "可引用角度: '当前最强LLM智能体在算法设计基准上只走了基线到最优差距的约17%';'大多数提交根本没有修改学习算法,只有少数修改的得分更高';'更多推理时间/努力更多只是让智能体更愿意修改训练算法'." Need "图书/PPT/简报" maybe "一句话结论" etc.
Need "原始材料" include "Fetched title", "Fetched description", "Abstract", "URL", "arXiv ID", authors, published. Need "英文关键词" maybe in "原始材料" section? Better include separate "英文关键词" after one sentence? User asked "保留英文标题、英文关键词和原始来源." Need maybe maintain "English title" as "AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement." Need perhaps "英文关键词: recursive self-improvement; LLM agents; algorithmic design; benchmark; training algorithms" from source. Need not say "keyword" if not explicitly in abstract? It's okay as "英文关键词(根据摘要提炼)"? User says 不要编造; if "keywords" not listed in source, we can say "英文关键词(根据摘要提炼)" to be transparent. Need maybe "关键词" from topics? Track foundation-model? But "英文关键词" not explicitly in source. We can derive from title and abstract. Need preserve "English title" in content. Need "原始来源" includes "URL: ..." "arXiv ID: ..." "Published: ..." "Updated: ..." "Primary category: cs.AI" "Categories: cs.AI, cs.CL, cs.LG" "Authors: ..." Good.
Need think about structure. Need first line exactly "# 知识卡片:AI4AI-Bench:评测LLM智能体在递归自我改进中的算法设计能力" That's a level-1 heading. Then maybe "> 英文标题:AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement" with blockquote? Fine.
Need "一句话结论" section. Maybe:
- 一句话结论:当前LLM智能体在“改写训练算法”这一RSI关键能力上仍很弱:AI4AI-Bench 的10个任务上,6个系统29种配置平均得分0.166,最强系统0.250,只走完从仓库自带算法到任务最优距离的不到1/5;大多数提交甚至完全没有改变模型的学习方式。
Need "事件概述或研究问题" section: "递归自我改进(RSI)问的是:一个AI系统能否改进产生AI系统的过程,使下一系统继承这一改进。论文认为该过程的核心是训练算法:更好的目标函数或更新规则,会改善每次后续运行的“算力-能力”兑换率,包括产生下一个智能体的那次运行。因此RSI是否可行,取决于智能体能否设计训练算法。现有基准无法孤立检验这种能力:已有评测往往靠收集数据或调超参取胜,且不能区分‘改变一次运行的执行方式’与‘改变模型的学习方式’。AI4AI-Bench 试图填补这一空白。" Good.
Need "方法/产品要点" section:
- "AI4AI-Bench 包含10个冻结的研究仓库,覆盖10个训练算法族。"
- "每个任务:智能体在单块B300上有4小时改写训练算法;随后其代码从零重跑,最多12小时。"
- "评分使用智能体不可见的固定评估器,并按与原仓库算法相同的流程比较。"
- "由于10个任务指标不可公度,所有任务都映射到同一尺度:0=无信息模型,0.1=仓库自带的算法,1.0=任务最优。"
- "发布内容包括任务套件、评估器和所有已打分提交,以便后续重复测量。" Need "frozen repositories" maybe "冻结" meaning locked to prevent leakage? We can say "冻结的研究仓库" from source.
Need "主要结果或产业意义" section:
- "平均分:0.166(6个系统、29种配置、全部10个任务);最佳系统0.250;最强系统也只完成从0.1到1.0距离的不到1/5。" Wait "best system reaches 0.250" "even the strongest closes under a fifth..." Need "best system" maybe not "strongest"? okay.
- "大多数提交从未改变模型的学习方式;少数改变学习方式的提交平均0.226,其余只有0.126。"
- "增加推理努力,主要买到的是‘愿意去改’:改变学习方式的提交比例从8%升至64%,平均分从0.094升至0.196。" Need "more reasoning effort" maybe "更多推理努力".
- "产业意义/评测意义:AI4AI-Bench 提供了一个可重复的基准,用于判断LLM智能体能否在训练算法层面实现递归自我改进;当前结果说明改进空间很大。" Need perhaps "产业意义" maybe "对AI4AI(AI for AI)方向有直接意义".
Need "为什么重要" section:
- "RSI 的可行性被归结为一个可度量问题:智能体能否设计训练算法。AI4AI-Bench 是试图隔离这一能力的基准。"
- "与已有基准的差异:已有基准可通过收集数据或调超参获胜,且无法分离‘执行方式改变’和‘学习方式改变’;AI4AI-Bench 用从零重跑的评估流程,将评估集中在训练算法本身。"
- "发布所有评估器与提交,使测量可重复,对后续系统变化开放。"
- "与既有脉络的关系:不同于PAST-Bench测个人智能体利用经验实现自我改进,也不同于自我精炼分类法侧重自我评估回路,AI4AI-Bench 直接测LLM智能体在算法设计上的能力,并把RSI的瓶颈定位到‘是否愿意/能够改变训练算法’这一层。"
Need "局限与不确定性" section:
- "摘要未说明10个训练算法族具体是什么、任务仓库来源、固定评估器的指标定义;需查阅全文(待核实)。"
- "摘要未披露6个系统/29种配置的模型名称与版本,也未说明推理努力如何量化和控制(待核实)。"
- "评分尺度的0.1为仓库自带算法,但0到1之间的分数未必是线性可解释的;跨任务映射的具体做法须以论文为准(待核实)。"
- "B300的具体产品定义/环境配置未在摘要中交代(待核实)。"
- "本文为arXiv预印本(v1),同行评议状态未说明(待核实)。"
Need maybe "用于图书/PPT/简报角度" section:
- "数据点:'AI4AI-Bench 上最强LLM智能体得分0.250,只完成从基线到最优差距的约17%'。"
- "现象:'是否修改学习算法'比'更多推理'更能区分得分;多数提交根本没触碰训练算法。"
- "启示:'更多推理努力主要让智能体愿意去修改训练算法,而不是直接保证更好结果'。"
- "可引用为:'递归自我改进的基准不应只看数据收集/超参调优,而应看算法设计本身'."
Need "原始材料" section:
- "英文标题:AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement"
- "英文关键词:recursive self-improvement; LLM agents; algorithmic design; benchmark; training algorithms" (based on abstract)
- "原始来源(URL):https://arxiv.org/abs/2608.20318v1"
- "arXiv ID: 2608.20318v1"
- "Authors: Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na"
- "Published: 2026-08-20T17:56:59Z; Updated: 2026-08-20T17:56:59Z"
- "Primary category: cs.AI; Categories: cs.AI, cs.CL, cs.LG" Need maybe "Fetched description" maybe not necessary if abstract included. But user requested "原始来源" not "all source material." Could include "摘要原文" maybe "Fetched description" enough? Need maybe "保留英文标题、英文关键词和原始来源" - include in raw section.
Need check no fabrication. "英文关键词" not explicitly in source; but label "根据摘要提炼" maybe. Need perhaps "英文关键词(根据摘要提炼)" to avoid premise. Good.
Need avoid "frontmatter" or "markdown" wrapper. The final answer should be Markdown text. Need not wrap in code block. Need no "" anywhere. Need maybe no extra outside.
Need think if "Track: academic, Topics: foundation-model, Manual title" from user should be included? They didn't ask to preserve. Could maybe include in "原始材料" no. But "Manual title" maybe "AI4AI-Bench..." We need not include track/topics.
Need maybe "已有相关卡片" with "PAST-Bench" "Frontis-MA1..." Need "与既有脉络的关系" should not duplicate "PAST-Bench" details. We can mention "在已有卡片中,PAST-Bench 关注个人智能体能否利用经验;本卡片关注的是能否设计训练算法" not repeat numbers. Good.
Need "一句话结论" perhaps "AI4AI-Bench 是衡量 RSI 中算法设计能力的基准;得分低。" Need maybe "recursive self-improvement" translation "递归自我改进". Need "算法设计" "训练算法" "目标函数或更新规则". Good.
Need be careful with "对 RSI 判定" "RSI 是否可行 therefore turns on whether an agent can design training algorithms" - We can say "论文主张" not fact. Need perhaps "作者提出" "论文认为". Use "论文认为" to attribute. Need "No benchmark isolates that ability" - "作者认为现有基准无法孤立检验这种能力" to attribute. Good.
Need perhaps "来源材料" "Fetched title" and "Fetched description" are given. Need not include "Fetched text excerpt" full abstract? Maybe "原始材料" section can include "摘要(英文原文)" from abstract. But user said "保留英文标题、英文关键词和原始来源" not necessarily full abstract. Yet "原始材料" could include "摘要(原文)" for reference. Need maybe not overlong but okay. Need "不要编造事实"; copying abstract is not fabrication. But maybe user asked "生成中文知识卡片草稿" and "中文正文"; source material in English okay. Need perhaps include "原始材料" with "Fetched description" not full abstract? We can include enough. Need maybe "英文摘要" not necessary. But "保留英文标题、英文关键词和原始来源" - "原始来源" could be URL and metadata. We can include "Abstract (English excerpt)" if helpful. But maybe too long. The desired oververbosity 5. Need moderate.
Need think if "AI4AI" name means "AI for AI"? Source not says. Could avoid expanding. "AI4AI-Bench" is title. Need not say "AI for AI". Actually n = "AI4AI" maybe "AI for AI" but not in abstract; don't fabricate. Could mention "AI4AI-Bench" as name.
Need maybe "10 training algorithm families" "10 frozen research repositories" - "spanning" not "covering" maybe. Need "frozen" could mean "固定不变" not "冻结" maybe "锁定版本". Use "冻结(frozen)" with English maybe.
Need "evaluator hidden from agent" - "固定评估器对智能体隐藏" important. Need "same procedure" "against repository's original algorithm under same procedure" - "在与原仓库算法相同的流程下,以仓库原始算法为对照评分." Good.
Need "run from scratch" - "从零重跑" should be "从头重新运行训练" to be clear. Need "code is then rerun from scratch" - "将智能体改写的代码从零开始重新运行(最多12小时)" Good.
Need "uninformative model" - "无信息模型" maybe "不含信息的模型" or "无信息量模型". Use "无信息模型" from source? Fine.
Need "compute-capability exchange rate" - "算力-能力兑换率" good.
Need "the next system inherits the improvement" - "下一系统继承该改进" Good.
Need "training algorithm" "a better objective or update rule" - "更好的目标函数或更新规则" Good.
Need "tells a change to how a run is executed apart from a change to how the model learns" - Need phrase "无法区分‘运行执行方式的改变’和‘模型学习方式的改变’" Good.
Need "closed under a fifth" maybe "不足五分之一" "接近五分之一" Actually "under a fifth" = "不到五分之一." Good.
Need "More reasoning effort mostly buys the willingness to go there" - "更多推理努力买来的主要是‘愿意去改’" a bit colloquial; perhaps "主要带来的是去修改训练算法的意愿". Good.
Need "less than a fifth of distance" "from 0.1 to 1.0" maybe if 0.25 -> 0.15 / 0.9 = 16.7%. Could state "约17%" but source says "under a fifth" not "about 17%" but derived. We can say "不到1/5" to be safe. Need not calculate 17%. But "under a fifth" okay.
Need "all 10 tasks" "29 configurations of 6 systems" maybe "6个系统的29种配置" Could be "6个系统`在全部10个任务上共29种配置" source says "Across 29 configurations of 6 systems on all 10 tasks" maybe "6 systems" not "6个系统"? We can say "6个系统(29种配置)在10个任务上的平均分" Good.
Need "evaluators hidden from agent" maybe "agent不可见" important because prevents gaming.
Need "with repository's original algorithm under same procedure" "against" - "与仓库原始算法在同一流程下比较". Good.
Need "the task suite, evaluators and every scored submission" "发布任务套件、评估器与所有已评分提交" Good.
Need maybe "no benchmark isolates that ability" "现有基准未能把这一能力单独测出来" Good.
Need perhaps "existing suites are won by collecting data or by tuning hyperparameters" "既有基准可以通过收集数据或调优超参数就获胜" Good.
Need "None tells a change to how a run is executed apart from a change to how the model learns." "无法把‘运行执行方式的变化’和‘模型学习方式的变化’分开" Good.
Need "RSI is feasible therefore turns on whether an agent can design training algorithms" "RSI是否可行,归结为智能体能否设计训练算法" Good.
Need "manual title" maybe "AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement" Use as English title.
Need maybe "事件概述或研究问题" include "研究问题" plus "event" maybe "Not an event; paper." Good.
Need "方法/产品要点" "产品" maybe "基准测试产品" Good.
Need "主要结果或产业意义" "产业意义" maybe "AI4AI-Bench released; can be used for model development." Good.
Need "为什么重要" "RSI feasibility" plus "incremental info in relation to existing cards" Good.
Need "局限与不确定性" maybe "待核实" list. Good.
Need "可用于图书/PPT/简报的角度" maybe "角度" list.
Need maybe "与既有脉络的关系" is not in requested "尽量使用如下小节" but optional. User says "尽量使用如下小节:" and "如果本条是对它们的延续或更新,请在“与既有脉络的关系”或“为什么重要”中简要说明增量信息." So include "与既有脉络的关系" maybe after "为什么重要". Good.
Need not include "frontmatter" so no "---". Good.
Need perhaps "英文关键词" should be under "原始材料" or separate. We'll include separate "英文关键词" maybe in "原始材料" section. Better:
## 原始材料
- 英文标题:AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
- 英文关键词(根据摘要提炼):recursive self-improvement; LLM agents; algorithmic design; benchmark; training algorithms
- 原始来源:https://arxiv.org/abs/2608.20318v1
...
Need "英文关键词" not hidden in raw. Good.
Need "保留英文标题" maybe in first line? The first line says Chinese title. Then maybe after first line, "英文标题:AI4AI-Bench..." Good. Need include "英文标题" in raw section too. Could include in "原始材料".
Need maybe "中文正文" means body Chinese, but English titles and keywords allowed. Good.
Need maybe "不要编造事实;无法从材料确认的部分写“待核实”." In "局限" we include "待核实" explicitly. Need maybe in "主要结果" not speculation. Good.
Need maybe "arXiv ID: 2608.20318v1" The date 2026-08-20 is future relative to current? Actually current date 2026-05-07, but source says published 2026-08-20; Wait current date is 2026-05-07, but source URL says 2608.20318? Wait 2608 maybe year 2026 month 08? It is future from current date. But user gave source, no issue. Need not mention date inconsistency. Hmm but "current date" is 2026-05-07 and source published 2026-08-20? Could be future; but maybe not relevant. Need maybe avoid "2026-08-20" if weird? But source material includes; user expects. Need include as metadata. No issue.
Need maybe "arXiv ID 2608.20318v1" Could be "2608" indicates 2026-08, but date 2026-08-20. Fine.
Need perhaps "track: academic, topics: foundation-model" optional? Could mention "类别:cs.AI; cs.CL; cs.LG" from source. Good.
Need think about "one B300" - If we say "单块 B300" no further. Need maybe "B300" might be "NVIDIA B300 GPU" but not in source; don't say NVIDIA. But in Chinese, "B300" might be ambiguous. "一块B300" okay.
Need perhaps "10 frozen research repositories" "10个冻结的研究仓库" Could "frozen" might be "the repositories are frozen" i.e., "锁定版研究代码仓库". Use "冻结(即固定版本)" perhaps. But "frozen" is source term; safe.
Need maybe "algorithmic design" "算法设计" in title. Good.
Need perhaps "compute-capability exchange rate" "算力-能力兑换率" maybe "算力-能力交换率" and "for every subsequent run" "每一次后续运行" good.
Need perhaps "role of reasoning effort" - "More reasoning effort mostly buys the willingness to go there" We can say "论文将更多推理努力的效果概括为:主要不是直接让改出来的算法更好,而是让智能体更愿意去改动模型的学习算法。" Is that supported? "mostly buys willingness to go there" yes. Need not add "not directly". Good.
Need perhaps "best system reaches 0.250" "最强系统" singular. Need "best system" maybe "其中最好的系统达到0.250". Good.
Need "evaluator hidden from agent" "固定评估器" maybe "评估器" in Chinese. Good.
Need "against repository's original algorithm under same procedure" "以仓库原始算法为基线比较" Good.
Need perhaps "The submissions show where that distance went" "提交结果说明差距去哪了" maybe "差距主要在于:大多数提交根本没有改变模型的学习方式" Good.
Need perhaps "release" "We release..." "作者开源了..." Good.
Need maybe "incommensurable metrics" "10个任务的指标不可公度" Good.
Need perhaps "0 is an uninformative model" "0=无信息模型" Maybe "无信息模型" might be unclear; "将0定义为不提供信息的模型" Good.
Need "score scale: 0.1 is algorithm repository ships" "0.1=仓库自带的原始算法" Good. "1.0 task optimum" "1.0=任务最优" Good.
Need perhaps "For every task, mapping onto one scale" "每个任务都映射到同一评分尺度" Good.
Need "frozen research repositories spanning 10 training algorithm families" "这10个仓库覆盖10个训练算法族" Good.
Need maybe "existing suites are won by collecting data or tuning hyperparameters" "现有基准的成绩可以通过收集数据或调超参获得" Good.
Need perhaps "all scored submissions" "所有已评分提交" Good.
Need "why important" include "if RSI feasible depends on whether agent can design training algorithm" Good.
Need perhaps "limitations" "not peer reviewed" maybe "arXiv预印本 v1;同行评议状态待核实" Could be considered "cannot confirm" yes. Need maybe "Publication date" from source but "v1" "updated same" okay.
Need maybe "no 'conclusion' heading"? We'll use sections.
Need ensure "第一行必须是一级标题" so first line exactly "# 知识卡片:AI4AI-Bench:评测LLM智能体在算法设计中的递归自我改进" Need choose title not too long? "知识卡片:AI4AI-Bench:..." The format says "# 知识卡片:中文标题" so okay. Need maybe use ":评测..." with colon? The title itself includes English and Chinese. Maybe "知识卡片:AI4AI-Bench——评测LLM智能体在递归自我改进中的算法设计能力" But "格式:# 知识卡片:中文标题" Maybe use "知识卡片:AI4AI-Bench:评测..." with two colons is okay. Need first line only. Good.
Need perhaps "中文标题" not "English title" first line. Good.
Need "保留英文标题、英文关键词和原始来源" in body. Good.
Need maybe "不要输出 frontmatter" - no frontmatter. Good.
Need maybe "不要把整段结果包裹在 ```markdown" - no code block. Good.
Need maybe "Markdown 正文" - use headings.
Need perhaps "Source material" includes "Fetched text excerpt" from source. Need maybe "原始材料" section with "摘要原文(英文)" if we include full abstract. But perhaps "Fetched description" is enough. However "原始来源" maybe URL and metadata. I'd include "Fetched description" as "摘要描述(英文)" in raw section, but not full "Abstract" to avoid too long. But user asks "根据以下来源材料生成" and "原始来源" maybe source material should be identifiable. We can include "原始来源(URL)" plus "arXiv metadata." Need maybe "Fetched title" not necessary if English title included.
Need maybe "英文关键词" "根据摘要提炼" maybe if not original. Good.
Need maybe "已有相关卡片" references "PAST-Bench" and "Frontis..." but not in output? User says "参考下面的已有相关卡片... if continuation..." We can include "与既有脉络的关系" section. Need avoid "copy" of existing card phrasing. Use new wording.
Need perhaps "AI4AI-Bench" vs "PAST-Bench" relation:
- "PAST-Bench 的切入点是‘经验保留/利用’:个人智能体能否从已积累经验获得后续任务改进。"
- "AI4AI-Bench 的切入点是‘训练算法设计’:LLM智能体能否改写训练算法本身,使下一系统继承改进。"
- "因此,AI4AI-Bench 更接近RSI的核心链路,而不是个人经验层面的改进;它同时提供了可重复的评估基准和公开提交。"
Need perhaps "Frontis-MA1/OpenMLE" from existing card no details; don't mention? Maybe "与Frontis-MA1/OpenMLE相关但不同" no need. Need maybe "有界自我精炼" from existing card: "更早的综述提出有界自我精炼与开放式RSI分类;AI4AI-Bench 是在‘算法设计’这一具体能力上做基准化,可视为对开放式RSI的可测量落地." Good.
Need perhaps not overstate "核心链路" maybe "论文认为 RSI 的可行性取决于算法设计" so yes.
Need perhaps "no benchmark isolates ability" "AI4AI-Bench 是第一个试图隔离该能力的基准" Did source say "No benchmark isolates that ability. We present..." It doesn't explicitly say "first", but "No benchmark isolates" and "We present" implies. Could say "论文提出 AI4AI-Bench 来隔离这一能力" rather than "first". Good.
Need perhaps "one B300" "on one B300" maybe "在单个 B300 环境" not "GPU". Good.
Need perhaps "agent has 4 hours ... to rewrite training algorithm; code then rerun from scratch for up to 12 hours" Need "总时间" maybe "4小时设计+12小时重训" Good.
Need perhaps "fixed evaluator hidden from the agent" "评估器固定且对智能体隐藏" good.
Need "scored by fixed evaluator hidden from the agent" "由智能体不可见的固定评估器打分" Good.
Need maybe "Because 10 metrics are incommensurable" "由于10个任务指标不可公度" good.
Need perhaps "mapped onto one scale" "统一映射到同一尺度" good.
Need "against repository's original algorithm under same procedure" "同一流程下以原仓库算法为对照" good.
Need maybe "willingness to go there" "去修改的意愿" good.
Need maybe "product要点" "产品" maybe not "product" but "benchmark product." Good.
Need maybe "fixed evaluator hidden from agent" "防止智能体针对评估器优化" This is an interpretation; source says hidden but not "prevent overfitting"; okay maybe not. Need no.
Need maybe "frozen repositories" "防止数据污染?" Could avoid explaining. Good.
Need perhaps "compute-capability exchange rate" "算力-能力兑换率" maybe "compute-capability exchange rate" in source. Good.
Need perhaps "recursive self-improvement" "RSI" maybe first use full.
Need perhaps "training algorithm families" "训练算法族" okay.
Need perhaps "task optimum" "任务最优" okay.
Need "one scale" "统一尺度" okay.
Need maybe "0.1 is algorithm repository ships" "0.1=仓库自带的算法" okay.
Need maybe "0 is an uninformative model" "0=无信息模型" maybe "uninformative model" might be "未从数据中获得信息的模型" or "无信息模型." Use "无信息模型" plus "(uninformative model)" perhaps.
Need perhaps "the algorithm that was already there" "已有算法" not "仓库 ships" Good.
Need maybe "most never change how the model learns at all" "大多数提交从头到尾都没有改变模型的学习方式" Good.
Need "the minority that do average 0.226 against 0.126 for rest" "少数真正改变学习方式的提交平均分0.226,其余只有0.126" Good.
Need "More reasoning effort mostly buys the willingness to go there, taking that minority from 8% of submissions to 64% and the mean score from 0.094 to 0.196." Need "takes minority from 8% of submissions to 64%" "将这类提交的占比从8%提升到64%" Good. "the mean score" "平均分从0.094到0.196" Good. Need "More reasoning effort" "更强的推理/更多推理努力" good.
Need perhaps "reasoning effort" can be "推理计算量" "reasoning effort" maybe "更多推理努力" not "推理能力". Good.
Need perhaps "best system reaches 0.250" "best system" "最好的系统" Good.
Need perhaps "mean score 0.166" "mean score" "平均分" Good.
Need perhaps "current LLM agents" "现有LLM智能体" Good.
Need perhaps "the strongest closes under a fifth of distance between algorithm that was already there and optimum" "即便最强系统也只完成了从仓库原始算法到最优解的不到1/5的距离" Good.
Need perhaps "existing suites are won by collecting data or by tuning hyperparameters" "现有基准擅长数据收集或超参调优就能得分" Good.
Need perhaps "none tells a change to how a run is executed apart from a change to how the model learns" "没有基准能区分‘运行方式改变’和‘模型学习方式改变’" Good.
Need perhaps "The process is the training algorithm" "该过程就是训练算法" Good.
Need perhaps "a better objective or update rule improves compute-capability exchange rate for every subsequent run, including the one that produces the next agent" "更好的目标函数或更新规则会提升每次后续运行的算力-能力兑换率,包括产生下一个智能体的那次运行" Good.
Need perhaps "RSI asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement" "递归自我改进(RSI)问的是:AI系统能否改进产生AI系统的过程,使下一系统继承改进?" Good.
Need perhaps "Therefore, whether RSI is feasible turns on whether an agent can design training algorithms" "因此RSI是否可行取决于智能体能否设计训练算法" Good.
Need perhaps "we release..." "作者发布..." Good.
Need perhaps "Peer review" "arXiv预印本" no.
Need perhaps "英文关键词" "recursive self-improvement; algorithmic design; LLM agents; benchmark; training algorithms" Could perhaps "algorithmic design" from title. Good.
Need perhaps "Manual title" in prompt "AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement" We should preserve exact english title. Good.
Need perhaps "来源材料" includes "Fetched description" exactly. Need perhaps "原始材料" section can include "Fetched description (English):" but maybe not needed. But user says "原始来源" maybe include "Source: arxiv abstract" not necessarily full. I'll include "原始材料" with "英文标题" etc. Need perhaps "Fetched title" no. Maybe "原始来源" includes URL and "Fetched description" from the prompt. Let's decide.
Let's draft in mind:
# 知识卡片:AI4AI-Bench:评测LLM智能体在递归自我改进中的算法设计能力
**英文标题**:AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
## 一句话结论
AI4AI-Bench 将递归自我改进(RSI)的关键问题落到一个可测量任务:LLM智能体能否改写训练算法本身。结果显示,6个系统的29种配置在10个任务上的平均分只有0.166,最佳系统0.250,即使最强系统也只完成从已有算法到任务最优距离的不到1/5;多数提交甚至从未改变模型的学习方式。
## 事件概述或研究问题
递归自我改进(RSI)问的是:一个AI系统能否改进“产生AI系统”的过程,使下一个系统继承该改进。论文认为,这一过程的核心是训练算法:更好的目标函数或更新规则,会提升每次后续运行的“算力-能力”兑换率,包括产生下一个智能体的那次运行。因此,RSI是否可行取决于智能体能否设计训练算法。
作者指出,现有基准无法单独测出这一能力:已有评测往往靠收集数据或调超参数就能获胜,而且无法把“改变一次运行的执行方式”和“改变模型的学习方式”区分开。AI4AI-Bench 的目标就是隔离并测量“智能体能否设计训练算法”。
## 方法/产品要点
- AI4AI-Bench 包含10个冻结的研究仓库,覆盖10个训练算法族。
- 每个任务中,智能体在单个B300上拥有4小时,重写该仓库的训练算法。
- 改写后的代码会从零开始重新运行,最多12小时;使用一个智能体不可见的固定评估器打分,并在同一流程下与仓库原始算法比较。
- 由于10个任务的指标不可公度,每个任务都被映射到同一评分尺度:0=无信息模型,0.1=仓库自带算法,1.0=任务最优。
- 作者发布任务套件、评估器以及所有已评分提交,以便后续对系统变化进行重复测量。
## 主要结果或产业意义
- 在全部10个任务、6个系统的29种配置上,平均分为0.166;最佳系统达到0.250。
- 即使最强系统,也只走完了从仓库原始算法(0.1)到任务最优(1.0)之间距离的不到1/5。
- 多数提交完全没改变模型的学习方式;少数改变学习方式的提交平均分0.226,其余只有0.126。
- 更多推理努力主要买到的是“愿意去改”:真正改变学习方式的提交占比从8%提升到64%,平均分从0.094提升到0.196。
- 对AI4AI方向的意义:AI4AI-Bench 提供了一个可重复的基准,用于衡量LLM智能体在训练算法层面的递归自我改进能力;当前水平离“有效RSI”仍有很大距离。
## 为什么重要
- 论文把RSI的可行性归结为“智能体能否设计训练算法”,AI4AI-Bench 是试图在此处建立可重复评测的基准。
- 与“通过收集数据或调超参就能获胜”的旧基准不同,它用“从零重跑+固定评估器+原算法对照”把分数集中在训练算法本身的改变上。
- 公开所有提交和评估器,意味着该测量可以随模型能力变化而重复,而不是一次性榜单。
## 与既有脉络的关系
- 与 PAST-Bench 关注“个人智能体能否从经验中获得后续改进”不同,AI4AI-Bench 关注的是“能否设计产生下一系统的训练算法”,更接近 RSI 的核心链路。
- 与“有界自我精炼/开放式RSI”分类法相比,AI4AI-Bench 提供了一个具体、可重复的基准化尝试,把“算法设计”作为开放式RSI的一个可测量入口。
## 局限与不确定性
- 摘要未列出10个训练算法族/研究仓库的具体名称,也未说明固定评估器的指标定义(待核实)。
- 摘要未披露6个系统、29种配置具体是哪些模型或提示配置(待核实)。
- “更多推理努力”如何量化、如何控制,摘要未说明(待核实)。
- B300的具体型号/环境配置未在摘要中交代(待核实)。
- 本文为 arXiv 预印本 v1;同行评议状态未在摘要中说明(待核实)。
## 可用于图书/PPT/简报的角度
- 一个可引用的数据锚点:“最强LLM智能体在AI4AI-Bench上只走完从基线算法到最优差距的不到1/5。”
- 一个值得讨论的现象:“多数提交根本没改训练算法;改了的系统平均分明显更高。”
- 一个提醒:“增加推理努力更主要地改变智能体‘去改算法’的意愿,而不保证改得更好。”
- 一个观点:“递归自我改进的评测,应聚焦于算法设计能力,而不能只看数据收集或超参调优。”
## 原始材料
- 英文标题:AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
- 英文关键词(根据摘要提炼):recursive self-improvement; LLM agents; algorithmic design; benchmark; training algorithms
- 原始来源:https://arxiv.org/abs/2608.20318v1
- arXiv ID:2608.20318v1
- 作者:Yizhe Chi, Wenyi Li, Deyao Hong, Xiaoqiu Wang, Mingju Gao, Kaisen Yang, Bingxiang He, Youjie Zheng, Calvin Xiao, Qinhuai Na
- 发布/更新:2026-08-20T17:56:59Z(v1,发布与更新相同)
- 分类:Primary category cs.AI; Categories cs.AI, cs.CL, cs.LG
Need check "每个任务中" - "In each task, an agent has 4 hours on one B300 to rewrite the training algorithm" Yes.
Need "固定评估器" "evaluator hidden from the agent" - maybe "对智能体隐藏的固定评估器" okay.
Need "从零开始重新运行" "rerun from scratch" yes.
Need "与仓库原始算法比较" "against repository's original algorithm under same procedure" yes.
Need "等待核实" in limitations. Good.
Need maybe "10个训练算法族/研究仓库" maybe "训练算法族" source says "10 training algorithm families" not "families / repositories" Good.
Need maybe "发布任务套件、评估器以及所有已评分提交" Good.
Need "6个系统的29种配置" Perhaps "Across 29 configurations of 6 systems on all 10 tasks" The phrase "29 configurations of 6 systems" might mean "29 configurations of systems" not "6 systems, 29 configurations." We can phrase "在全部10个任务上,6个系统的29种配置" exactly from source. Good.
Need perhaps "最佳系统达到0.250" "the best system reaches 0.250" yes.
Need "多数提交完全没改变模型的学习方式" "most never change how the model learns at all" yes.
Need "少数改变学习方式的提交平均分0.226,其余只有0.126" "the minority that do average 0.226 against 0.126 for the rest" yes.
Need "更多推理努力主要买到的是‘愿意去改’" "More reasoning effort mostly buys the willingness to go there" yes. Need maybe "买" colloquial; maybe "主要带来的是" better.
Need "真正改变学习方式的提交占比从8%提升到64%" "taking that minority from 8% of submissions to 64%" yes. Need "more reasoning effort" "更强推理努力" maybe "more reasoning effort" might be "更多的推理努力(例如更长的思考时间/推理计算)" But source doesn't specify. Maybe "更多推理努力" okay.
Need "当前水平离‘有效RSI’仍有很大距离" Is this conclusion from scores yes. Need maybe "有效RSI" not in source; "RSI feasible" yes. Could say "仍有很大距离" okay.
Need "与既有脉络" perhaps "PAST-Bench 关注经验" Good. Need avoid "core chain" maybe okay. Need perhaps "Frontis-MA1/OpenMLE" not included. Good.
Need "原始材料" Need "Fetched description" maybe maybe "原始材料" should include "摘要原文" because "原始来源" should be traceable. Could include "摘要原文(英文)" after metadata? But maybe too long. Need perhaps include "原始来源" enough. However user asks "保留英文标题、英文关键词和原始来源" not "摘要原文." Fine.
Need "英文关键词" "根据摘要提炼" maybe okay. But "recursive self-improvement" appears in abstract; "algorithmic design" in title; "LLM agents" in title; "benchmark" in title; "training algorithms" in abstract. Good.
Need perhaps "来源材料" has "Fetched description: Recursive self-improvement..." We can maybe include "Fetched description" as "原始材料" maybe:
- "Fetched description(英文摘要):Recursive self-improvement (RSI) asks..." Hmm if include full fetched description, it duplicates abstract. Not necessary.
Need perhaps "已提供来源" "URL" "arXiv ID" "authors" "published" etc. Good.
Need "不要编造事实;无法从材料确认的部分写“待核实”." Need "作者发布" in source "We release" yes. Good.
Need perhaps "B300" "单个B300" no "GPU" no fabrication. Good.
Need perhaps "10个冻结的研究仓库" "冻结" maybe "frozen" could mean "版本锁定" but no explanation. Fine.
Need perhaps "摘要未列出..." Need "待核实" maybe "需要阅读全文核实" rather than "待核实" exactly. User asked "写“待核实”"; include the exact phrase. Good.
Need perhaps "incommensurable" "不可公度" maybe a bit technical; perhaps "无法直接比较的度量" maybe "指标不可公度(量纲/含义不同)" Good. Need "公度" may be obscure. Could say "由于10个任务各自的指标不具可比性" Good. But "不可公度" is accurate. Need maybe "同一尺度" enough.
Need perhaps "task optimum" "任务最优" okay.
Need perhaps "0=无信息模型" "uninformative model" maybe "无信息量的模型" better. Use "无信息量的模型" to make Chinese clearer. But source says "uninformative model"; "无信息量模型" okay. Let's use "无信息量模型(uninformative model)" perhaps.
Need perhaps "score scale" "0=无信息量模型,0.1=仓库自带算法,1.0=任务最优" Good.
Need perhaps "compute-capability exchange rate" "算力-能力兑换率" perhaps "算力与能力的兑换率" Good.
Need perhaps "training algorithm family" "训练算法族" Good.
Need perhaps "event overview" "研究问题" could include "No benchmark isolates..." Good.
Need perhaps "why important" "open-source" maybe "发布任务套件..." repeated from methods. But okay. Need perhaps "为什么重要" not repeat too much. Could mention "增量信息" in "与既有脉络的关系" section.
Need perhaps "limitations" "not enough to know if RSI feasible" maybe "摘要没有说明AI4AI-Bench 对RSI可行性的最终判定;当前只是测量智能体算法设计能力." Good.
Need perhaps "with existing cards" "PAST-Bench" no duplicate facts? We mention PAST-Bench conceptually. Good.
Need perhaps "frontmatter" no.
Need perhaps "不要将整段结果包裹在代码块中" no code block.
Need perhaps "first line must be level 1 heading" yes.
Need perhaps "Chinese title" "知识卡片:AI4AI-Bench:评测LLM智能体在递归自我改进中的算法设计能力" Good.
Need perhaps "Source material" "Fetched title" in original "Fetched title: AI4AI-Bench..." We preserve. Good.
Need maybe "原文标题" maybe "英文标题" in raw. Good.
Need "英文关键词" exact "recursive self-improvement; LLM agents; algorithmic design; benchmark; training algorithms" Good.
Need perhaps "Track: academic" no need.
Need maybe "Topics: foundation-model" no need.
Need maybe "Manual title" no need.
Need perhaps "AI4AI" in title "AI4AI-Bench" not "AI4AI-Bench" maybe okay.
Need maybe "Mingju Gao" etc. Good.
Need perhaps "v1" "Published 2026-08-20T17:56:59Z" "Updated same" Good.
Need perhaps "原始材料" with "Primary category cs.AI; Categories cs.AI, cs.CL, cs.LG" Good.
Need perhaps "limitations" "10个任务的具体难度" etc. Need "摘要没有交代..." Good.
Need maybe "reasoning effort" "more reasoning effort" perhaps "更强的推理努力" not "推理努力" because source says "more reasoning effort" not "stronger". Use "更多推理努力" consistently.
Need perhaps "mean score from 0.094 to 0.196" "平均分从0.094升至0.196" Good.
Need perhaps "minority from 8% to 64%" "that minority" refers "submissions that change how model learns" yes. Good.
Need perhaps "proportion" "占比" Good.
Need maybe "closes under a fifth" "closes" maybe "缩小差距" - "只缩小了不到1/5" Good. Need "distance between algorithm that was already there and optimum" "从已有算法到最优的距离" Good.
Need perhaps "Best system reaches 0.250" "best system" maybe "其中最好系统" Good.
Need maybe "all 10 tasks" "全部10个任务" Good.
Need maybe "evaluator hidden from agent" "hidden" important to avoid "fixed evaluator" "固定评估器" Good.
Need perhaps "score by fixed evaluator hidden from the agent" "由固定评估器打分,且该评估器对智能体隐藏" Good.
Need perhaps "same procedure" "同一流程" yes.
Need perhaps "the repository's original algorithm under the same procedure" "以仓库原始算法为对照,使用相同流程" Good.
Need perhaps "existing suites are won by collecting data or tuning hyperparameters" "现有基准往往靠收集数据或调超参数就能获胜" Good.
Need perhaps "None tells a change to how a run is executed apart from a change to how the model learns" "没有一个基准能把运行执行方式的改变与模型学习方式的改变区分开" Good.
Need perhaps "whether an agent can design training algorithms" "智能体能否设计训练算法" Good.
Need perhaps "that process is the training algorithm" "该过程就是训练算法" Good.
Need perhaps "compute-capability exchange rate" "算力-能力兑换率" Good.
Need perhaps "for every subsequent run, including the one that produces the next agent" "每一次后续运行,包括产生下一个智能体的那一次" Good.
Need perhaps "recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inherits the improvement." "递归自我改进(RSI)问的是:AI系统能否改进产生AI系统的过程,使下一系统继承这一改进?" Good.
Need perhaps "中文正文" all sections Chinese. Good.
Need maybe "one B300" "单块B300" rather than "单个B300" because "one B300" could be "一块" but not "GPU". Use "单个B300" neutral. In Chinese "单个B300计算环境" maybe too much. Use "单块B300" but "块" implies physical unit, maybe okay. Source "one B300" likely one accelerator. "单块B300" is okay. Could say "一个B300(原文为 one B300)" to avoid interpretation. Maybe "一个B300环境" weird. Let's use "单个B300" (not "块") to avoid physical assumption. "智能体在一个B300上有4小时" Good.
Need perhaps "frozen" "冻结" maybe "固定不可变" but "frozen research repositories" means "研究仓库被冻结" to prevent modifications? Good.
Need perhaps "task suite" "任务套件" okay.
Need perhaps "original algorithm" "仓库原始算法" Good.
Need perhaps "the evaluation can be repeated as systems change" "当这些系统变化时,测量可重复" Good.
Need perhaps "for up to 12 hours" "最多12小时" Good.
Need perhaps "uninformative model" "无信息量模型" Good.
Need perhaps "incommensurable" "不可公度" maybe "不可通约" but "不可公度" okay. Need maybe "因为10个任务的评价指标各不相同、无法直接换算" Good. Let's perhaps phrase in "方法/产品要点": "由于10个任务各自的指标不具可比性(incommensurable)" Good. Then scale.
Need perhaps "task optimum" "任务最优值" Good.
Need perhaps "with repository's original algorithm under the same procedure" "在与仓库原始算法相同的重跑流程下进行对照" Good.
Need perhaps "more reasoning effort mostly buys willingness to go there" "更多推理努力主要买到的是‘愿意去修改’" maybe "buy" english phrase; use "带来的主要是……意愿" Good.
Need perhaps "limitations" "摘要没有说明..." Need "待核实" in each bullet maybe. Use "(待核实)" at end of each bullet. Good.
Need perhaps "not peer-reviewed" "arXiv预印本" maybe source doesn't say "preprint" but URL arxiv.org. Fine.
Need perhaps "source material" "Fetched description" maybe if include "Fetched description" can be seen as source. Need perhaps "原始材料" section:
- 原始来源:https://arxiv.org/abs/2608.20318v1
- 英文标题:...
- 英文关键词:...
- 摘要(原文):Recursive self-improvement (RSI) asks whether an AI system can improve the process...
Need include full abstract? The user gave "Fetched text excerpt" containing abstract. "原始材料" can have "摘要原文" to preserve source. But "不要编造事实" and "保留英文..." Maybe okay. However too long? Not too long. But maybe "不要 wrap code block" no issue. Let's consider adding "原始材料" with "摘要(英文原文)" maybe a blockquote? It would be English; but user asked Chinese正文 not necessarily no English. "保留英文标题/关键词/原始来源" okay. But "中文正文" maybe body should be Chinese; raw abstract is okay. But maybe not needed. To keep concise, not include full abstract. But "原始材料" might be richer if include "Fetched description" exactly. We can include "原始材料" with "Fetched description" bullet? Hmm.
The phrase "Source material:" in user prompt includes "Fetched title" "Fetched description" etc. They want "原始来源" preserved, not "copy all source." We can include "原始材料" with URL and metadata. Good.
Need perhaps "Manual title" maybe "AI4AI-Bench..." already.
Need perhaps "arXiv ID 2608.20318v1" with "v1" okay.
Need perhaps "Published" "Updated" in raw. Good.
Need perhaps "Primary category" "cs.AI" and categories. Good.
Need perhaps "Track: academic" not necessary. But maybe "类别" from source "cs.AI, cs.CL, cs.LG" yes.
Need perhaps "原创性" no.
Need perhaps "one sentence conclusion" maybe "一句话结论" exactly.
Need perhaps "事件概述或研究问题" section heading maybe "事件概述/研究问题" but user says "事件概述或研究问题"; use "事件概述或研究问题" exactly. Good.
Need perhaps "方法/产品要点" exactly. Good.
Need perhaps "主要结果或产业意义" exactly. Good.
Need perhaps "为什么重要" exactly. Good.
Need perhaps "局限与不确定性" exactly. Good.
Need perhaps "可用于图书/PPT/简报的角度" exactly. Good.
Need perhaps "原始材料" exactly. Good.
Need "与既有脉络的关系" heading exactly maybe "与既有脉络的关系". Good.
Need perhaps "一句话结论" "AI4AI-Bench 是..." Good.
Need perhaps "为什么要避免重复" "参考已有相关卡片,避免与它们重复陈述同一事实" We already.
Need perhaps "if this is continuation/update" "与既有脉络" includes.
Need maybe "已有相关卡片" "PAST-Bench" focuses personal agents. "AI4AI-Bench" maybe not "recursive self-improvement" in the sense of agent improving its own output? Actually source says "recursive self-improvement" asks if AI system can improve process that produces AI systems; training algorithm. Existing card about PAST-Bench "personal agents recursive self-improvement" maybe "个人智能体从经验中改进" not "training algorithm." Need mention "不是经验层面,而是训练算法层面" good.
Need perhaps "Frontis-MA1/OpenMLE" "AI4AI recursive self-improvement system" no detail. Could maybe mention "与Frontis-MA1/OpenMLE这类系统报告不同,AI4AI-Bench是基准而不是系统" but not needed. Maybe in "与既有脉络" mention "与既有系统卡片不同,它提供的是评估基准而非新的自我改进系统." Is that true? Existing Frontis-MA1/OpenMLE are "systems" maybe yes. But source doesn't mention them. This is from existing card title. Could say "与已有AI4AI系统卡片(如Frontis-MA1/OpenMLE)相比,AI4AI-Bench 不是提出新的自我改进系统,而是提出可重复测量算法设计能力的基准。" This is a valid incremental relation. Should we include? Maybe yes. But user only provided "已有相关卡片" with no details; "Frontis-MA1与OpenMLE" title says "面向机器学习工程的AI4AI递归自我改进系统" no one-sentence conclusion. We can mention carefully "相较已有的AI4AI系统型卡片,本卡片的增量是..." Good.
Need perhaps "PAST-Bench" and "Frontis-MA1/OpenMLE" maybe all from "已有相关卡片"; include relation: "与 PAST-Bench 从经验保留角度测个人智能体不同;与Frontis-MA1/OpenMLE这类系统报告不同,AI4AI-Bench 提供的是基准测试而不是新系统." Good. Need not duplicate exact facts.
Need perhaps "Frontis-MA1" is "AI4AI recursive self-improvement system" but no details. We can say "它不提出新系统,而是提出评测基准." Good.
Need perhaps "One sentence conclusion" "现有LLM智能体在算法设计上表现有限" Good.
Need perhaps "来源材料" "arXiv abstract" yes.
Need maybe "B300" "单块B300" no.
Need perhaps "hidden evaluator" "固定评估器对智能体隐藏" Good.
Need perhaps "Task evaluation" "score against original algorithm" Good.
Need perhaps "0.1 is algorithm repository ships" "0.1=仓库自带算法" Good.
Need perhaps "original repository" maybe "仓库自带的算法" Good.
Need perhaps "uninformative model" "无信息量模型" Good.
Need perhaps "reasoning effort" "more reasoning effort" maybe "更多推理努力主要增加的是智能体去改动学习算法的倾向" Good.
Need perhaps "the best system reaches 0.250" "even the strongest closes under a fifth..." Need "best system" and "strongest" are same? Source "best system reaches 0.250: even the strongest closes..." likely yes. Use "最佳/最强系统" but "best system" and "strongest" perhaps same. In one sentence "最佳系统达到0.250;即使最强系统也只缩小不到1/5差距" okay.
Need perhaps "mean score 0.166" "Across 29 configurations of 6 systems on all 10 tasks" maybe "6个系统的29种配置" Could be "29 configurations of 6 systems" = 6 systems, 29 configurations total? We can write "6个系统的29种配置" okay.
Need perhaps "29 configurations of 6 systems" could mean "29 configurations [of models] from 6 systems" not "6 systems and 29 configs." We can say "29种系统配置(涉及6个系统)" Good: "在全部10个任务上,29种系统配置(涉及6个系统)的平均分..." Actually "29 configurations of 6 systems" maybe "6 systems, 29 configurations" yes. Good.
Need perhaps "B300" "one B300" maybe "一个B300" not "single B300" okay.
Need perhaps "ML engineering" no.
Need perhaps "not fabricate" "AI4AI-Bench 是第一个..." not in source; avoid "first". Good.
Need perhaps "existing suites are won..." "现有基准可以通过收集数据或调整超参数获胜" Good.
Need perhaps "有界自我精炼与开放式RSI" from existing card "分类法" not "existing card" maybe "论文" but source not mention. We can mention "在已有脉络中,一篇2026年的综述提出了..." But perhaps not necessary? The "已有相关卡片" includes "AI中的递归自我改进:从有界自我精炼到自主研究循环" with categorization. We can say "与此前综述对有界自我精炼/开放式RSI的分类法相比..." Good.
Need perhaps "AI4AI-Bench 作为开放式RSI的可测量入口" Maybe "开放式RSI" not in source; but in existing card. Need maybe "可视为..." is interpretation, maybe okay but "不要编造"; maybe "可视为" is not fact but analytic. To be safe, phrase "从既有脉络看,它把'算法设计'这一能力变成了可重复测量的基准." Good.
Need perhaps "待核实" in "with existing" no.
Need perhaps "available for book/PPT/briefing" "数据锚点" Good.
Need perhaps "Final answer" no extra. Need produce final.
Need maybe "第一行" "一级标题" exactly no blank before? Yes.
Need maybe "not output frontmatter" no.
Need perhaps "不要将整段结果包裹在 ```markdown" no code block. Good.
Need perhaps "Markdown" use #, ##, bullets, bold.
Need perhaps "A one-sentence conclusion" maybe use "一句话结论:" but "一句话结论" as section perhaps. User says "尽量使用如下小节:一句话结论..." Could be ## 一句话结论 or bold. Better use ## 一句话结论 to align. Need "英文标题" perhaps before "一句话结论"? First line heading, then maybe a line with English title. Should we put "英文标题..." after first line before first section. Good.
Need perhaps "第一行 must be level 1 title" "第一行" includes no blank before; okay. Good.
Need perhaps "中文标题" "知识卡片:AI4AI-Bench:评测LLM智能体在递归自我改进中的算法设计能力" The colon after "AI4AI-Bench" might be inside title. Good.
Need perhaps "原始材料" "原始来源" "URL" and "arXiv ID" etc. Good.
Let's think about any possible issue with "arXiv ID 2608.20318v1" - The arxiv ID format "2608.20318" seems future; but original source says. Fine.
Need perhaps "Published 2026-08-20" and "Updated 2026-08-20" same. Good.
Need perhaps "Primary category cs.AI; Categories cs.AI, cs.CL, cs.LG" Good.
Need perhaps "one B300" in methods "4小时" "最多12小时" "single B300" yes.
Need perhaps "score scale" "0=无信息量模型" maybe "uninformative model" perhaps "无信息模型" not "无信息量" okay. Use "无信息量模型" to avoid "uninformative" confusion.
Need perhaps "In each task" "each of the 10 tasks" Good.
Need perhaps "the task suite, evaluators and every scored submission" "所有已评分提交" Good.
Need perhaps "Industry significance" "产业意义" perhaps "对AI4AI方向" Good. Need "AI4AI" maybe no expansion. Good.
Need perhaps "More reasoning effort mostly buys willingness" "mostly" "主要" Good.
Need perhaps "submissions show where that distance went" "差距去哪儿了" maybe in results. Good.
Need perhaps "closes under a fifth" "只缩小不到1/5" Good.
Need perhaps "existing suites are won by collecting data or by tuning hyperparameters" "旧基准" "现有" Good.
Need perhaps "No benchmark isolates that ability" "无法单独测出" Good.
Need perhaps "none tells a change to how a run is executed apart from a change to how the model learns" "不能区分运行执行方式与模型学习方式的变化" Good.
Need perhaps "we release" "作者开源" Good.
Need perhaps "fixed evaluator hidden from agent" "固定评估器对智能体隐藏" Good.
Need perhaps "same procedure" "相同流程" Good.
Need perhaps "The repository's original algorithm" "仓库原始算法" Good.
Need perhaps "incommensurable" "不可公度" maybe "不可直接比较" Good.
Need perhaps "task optimum" "任务最优值" Good.
Need perhaps "one B300" "B300" maybe "单块B300" okay.