{"id":1696,"date":"2026-09-08T09:03:47","date_gmt":"2026-09-08T01:03:47","guid":{"rendered":"https:\/\/blog.liu-qi.cn\/2026\/09\/08\/x-daily-2026-09-07\/"},"modified":"2026-09-08T09:03:47","modified_gmt":"2026-09-08T01:03:47","slug":"x-daily-2026-09-07","status":"publish","type":"post","link":"https:\/\/en.blog.liu-qi.cn\/2026\/09\/08\/x-daily-2026-09-07\/","title":{"rendered":"X Platform September 7 AI Brief | OpenAI's Automation Research Milestone, GPT-6\/Astra's Continuous Task Capability Expansion, and Programming AI Covers Formal Mathematics"},"content":{"rendered":"<h2 id=\"topic-c185e6f5d4\">OpenAI Discloses Milestone in Automation Research, But Recursive Self-Improvement Remains Constrained by Safety Boundaries<\/h2>\n<p>The core change evident in yesterday&#8217;s discussion is that OpenAI&#8217;s public materials have advanced &#8220;AI participating in AI R&amp;D&#8221; from a concept to an internal engineering milestone: a blogger relayed that the team has already implemented an &#8220;automated research intern&#8221; that, guided by human-set goals and instructions, can independently complete tasks that typically require skilled researchers several days to handle. The next target is set for an automated AI researcher by March 2028. The materials also mention that internal Agent usage has reached a significant scale, but these figures and timelines are from public materials and blogger accounts and have not been independently verified this round.<\/p>\n<p>Notably, the same information chain does not present the capability progress as &#8220;solved RSI.&#8221; OpenAI&#8217;s Chief Scientist&#8217;s public article emphasizes that model capability improvements may outpace monitoring capabilities, and the method of observing intent solely through chain-of-thought is becoming unreliable; whether to continue accelerating still depends on whether humans can maintain control and establish common safety thresholds.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@merettm: <a href=\"https:\/\/x.com\/merettm\/status\/2096630018495377464\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/merettm\/status\/2096630018495377464<\/a><\/li>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2096636734838841827\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2096636734838841827<\/a><\/li>\n<li>@xiaohu: <a href=\"https:\/\/x.com\/xiaohu\/status\/2096790776110067810\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/xiaohu\/status\/2096790776110067810<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-735e3811c5\">GPT-6\/Astra&#8217;s Computer Use Has Progressed from Single Demonstrations to Continuous Tasks<\/h2>\n<p>Multiple bloggers have demonstrated GPT-6\/Astra performing continuous operations in real interfaces: one recorded it clearing all 48 levels of an &#8220;I&#8217;m not a robot&#8221; game, while another claimed it ran for about 6 hours in Minecraft to complete the chained task of mining diamonds, which involves gathering resources, crafting tools, and exploration. In image workflows, a blogger also showed Astra creating a drawing from a sketch to a high degree of fidelity in Photoshop, and using Suno to read recommended song prompts, batch organize them, and generate a new song.<\/p>\n<p>These are visible demonstrations, practical tests, and accounts from bloggers or users, indicating that the capability boundary is expanding from &#8220;generating content&#8221; to &#8220;understanding interfaces and executing multi-step tasks&#8221;; they are not equivalent to official performance promises for all users across all software environments.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@xiaohu: <a href=\"https:\/\/x.com\/xiaohu\/status\/2096949627966931186\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/xiaohu\/status\/2096949627966931186<\/a><\/li>\n<li>@ZHO_ZHO_ZHO: <a href=\"https:\/\/x.com\/ZHO_ZHO_ZHO\/status\/2096900358169894916\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/ZHO_ZHO_ZHO\/status\/2096900358169894916<\/a><\/li>\n<li>@vista8: <a href=\"https:\/\/x.com\/vista8\/status\/2096768585956024379\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/vista8\/status\/2096768585956024379<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-3b9afe40ca\">Programming-Capable AI Begins to Cover Formal Mathematics, But Task Orchestration Remains a Bottleneck<\/h2>\n<p>Regarding the case involving Replit founder Amjad Masad, a blogger relayed that Claude generated approximately 13 million lines of Lean code over 11 days to formalize the existing proof of Fermat&#8217;s Last Theorem into a program that can be checked step-by-step. The value of this case lies not in the &#8220;lines of code&#8221; itself, but in extending AI programming from website and app prototypes to programmable research tasks like mathematical proofs and experimental verification.<\/p>\n<p>The same account also retains the constraints: the project experienced failures and only progressed successfully after adding task dependency and collaborative progress management. Therefore, the judgment that &#8220;almost anything that can be turned into a programming problem can be solved&#8221; belongs to the participants; a more cautious current conclusion is that key bottlenecks for long-cycle research tasks still include decomposition, dependency management, and process supervision.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2096988178834129400\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2096988178834129400<\/a><\/li>\n<li>@amasad: <a href=\"https:\/\/x.com\/amasad\/status\/2096936109817135331\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/amasad\/status\/2096936109817135331<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-4560084068\">Codex&#8217;s New Context Mechanism Shows a Clear Cost-Benefit Trade-off<\/h2>\n<p>MaxForAI provided a &#8220;post-sale&#8221; correction to the Codex new context feature recommended a few days ago: after further practical testing, he believes this experimental feature can indeed significantly reduce token consumption, but as the conversation rounds lengthen, the model forgets previous requirements, confuses sessions, and even exhibits a noticeable decline in quality. He speculates the cause might be incomplete note-taking or failed historical retrieval, but the specific link cannot be distinguished for now.<\/p>\n<p>The reusable conclusion from this experience is: for long tasks, one cannot only consider the cost after context compression; one must also observe error correction, re-explaining context, and result stability in the latter half of the task. Until the official version is more stable, a trade-off remains between saving tokens and usability.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2096797970582921595\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2096797970582921595<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-d94d2acea5\">Competitive Focus of Computer-use Tools Shifts to Background Control, Learnability, and Transferability<\/h2>\n<p>The Huashu tool introduced by AlchainHust focuses on &#8220;completing tasks&#8221; rather than simulating clicks: in their test, a Feishu multi-dimensional table task that initially required over 200 steps of trial and error was reduced to about 10 steps after experience was accumulated. The tool also selects control methods based on website architecture and tries not to interfere with the user&#8217;s foreground operations. The design of Huashu-mac-use also emphasizes background execution, choosing between code or interface operations based on software architecture, and accumulating trial-and-error experience back into skills.<\/p>\n<p>Another practical test shows that GPT-6\/Astra&#8217;s Computer Use and Browser Use allowed a user to complete the entire development and publishing process for two iOS Apps and one Obsidian plugin in one go. What is more noteworthy now is not a single instance of &#8220;can it click,&#8221; but whether skills can be learned, reused, and transferred across different software; the aforementioned effects still primarily come from developer self-reports and individual workflow cases.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@AlchainHust: <a href=\"https:\/\/x.com\/AlchainHust\/status\/2096968390409949488\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/AlchainHust\/status\/2096968390409949488<\/a><\/li>\n<li>@AlchainHust: <a href=\"https:\/\/x.com\/AlchainHust\/status\/2096845689653543029\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/AlchainHust\/status\/2096845689653543029<\/a><\/li>\n<li>@vista8: <a href=\"https:\/\/x.com\/vista8\/status\/2096908619828871659\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/vista8\/status\/2096908619828871659<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-ed50bbc0ed\">Bilibili&#8217;s Build in Public Event Demonstrates a User Co-Creation Loop for AI Products<\/h2>\n<p>When reviewing the offline awards ceremony for the Bilibili BIP event in Shanghai, MaxForAI mentioned that the first-prize winner, &#8220;Neko Project Neko&#8221; with a 1 million yuan prize, started as just an AI companion prototype. It later expanded to include game plugins for titles like Minecraft, Red Alert, and Mahjong Soul through live-streamed trials, viewer feedback, and GitHub contributions. Another creator placed their game on Bilibili&#8217;s Toy platform, allowing viewers to try it directly after watching a video, report bugs, and suggest features.<\/p>\n<p>The core insight from such cases is not the event&#8217;s popularity, but the cycle formed between content, product, and user feedback: AI lowers the barrier to prototype development, publicizing the development process attracts testers and collaborators, and vertical users then supplement needs the team itself was unaware of. External reports also stated the event had 13,400 participants, but this figure is cited here only as a claim found in visible posts.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2096975148595491204\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2096975148595491204<\/a><\/li>\n<li>@hey_madni: <a href=\"https:\/\/x.com\/hey_madni\/status\/2096863698791137286\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/hey_madni\/status\/2096863698791137286<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-1374d722ad\">GPT-6\/Astra&#8217;s Compute Scale Intensifies Domestic Catch-Up Pressure<\/h2>\n<p>MaxForAI cited Jensen Huang&#8217;s public post stating that GPT-6 Astra was trained using approximately 100,000 Grace Blackwell NVLink72 units, with plans to deploy around 400,000 more subsequently. Based on this, he described the pressure felt by China&#8217;s algorithm and academic circles regarding the compute gap, while also noting that the online traffic of models hosted on domestic chips from HiSilicon, T-Head, and others remains a visible local advancement.<\/p>\n<p>It is necessary to distinguish the levels of fact here: the figures of 100,000 and 400,000 units are from Jensen Huang&#8217;s public post. The &#8220;catch-up pressure&#8221; and progress of domestic chips are the blogger&#8217;s interpretation. This single citation should not be presented as independently verified industry statistics. It is better suited as a visible signal within the discussion space of frontier model competition, where training scale and supply chain capabilities continue to widen.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@JensenHuang: <a href=\"https:\/\/x.com\/JensenHuang\/status\/2096700264569090384\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/JensenHuang\/status\/2096700264569090384<\/a><\/li>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2096831318642635224\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2096831318642635224<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-328bc3addc\">Anthropic&#8217;s $1.6 Trillion Valuation is a Model Projection, Not a Confirmed Valuation<\/h2>\n<p>MaxForAI cited @jonbma&#8217;s calculation: Starting from the previously reported annualized revenue run rate for Anthropic exceeding $65 billion as of late July, and assuming an additional $10 billion in ARR added each month thereafter, the model projects a total ARR of approximately $1.15 trillion by year-end. After deducting revenue shares for Meta, cloud platforms, and revenue deemed related to distillation with Chinese labs, the net ARR is about $820 billion. Applying a 20x EV\/ARR multiple yields an enterprise value of approximately $1.64 trillion.<\/p>\n<p>This is a scenario analysis dependent on revenue growth, deductions, and valuation multiples, and should not be equated with Anthropic having completed a funding round or a market-confirmed valuation. The confirmable information is merely the appearance of this calculation on social media and the corresponding PreStocks price discussion.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@jonbma: <a href=\"https:\/\/x.com\/jonbma\/status\/2096701856189714634\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/jonbma\/status\/2096701856189714634<\/a><\/li>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2096903249760858223\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2096903249760858223<\/a><\/li>\n<\/ul>\n<p>Statistics: Timeline Scanned Posts=720 Matching Bloggers=62 Total Matching Tweets=369 Weighted Tweet Score=302.25 Original Tweets=161 RT Tweets=52 Crawl Attempts=5 Boundary Coverage Status=tail_confidently_crossed_target_boundary<\/p>\n","protected":false},"excerpt":{"rendered":"<p>AI R&#038;D automation achieves engineering progress, model capabilities for executing complex tasks expand, and programming AI begins handling research tasks like mathematical proofs.<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[19],"class_list":["post-1696","post","type-post","status-publish","format-standard","hentry","category-brief","tag-x--ai-"],"_links":{"self":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts\/1696","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/comments?post=1696"}],"version-history":[{"count":0,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts\/1696\/revisions"}],"wp:attachment":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/media?parent=1696"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/categories?post=1696"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/tags?post=1696"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}