{"id":1678,"date":"2026-09-03T09:04:50","date_gmt":"2026-09-03T01:04:50","guid":{"rendered":"https:\/\/blog.liu-qi.cn\/2026\/09\/03\/x-daily-2026-09-02\/"},"modified":"2026-09-03T09:04:50","modified_gmt":"2026-09-03T01:04:50","slug":"x-daily-2026-09-02","status":"publish","type":"post","link":"https:\/\/en.blog.liu-qi.cn\/2026\/09\/03\/x-daily-2026-09-02\/","title":{"rendered":"X Platform September 2 AI Brief | Claude Fable 5.1 Significantly Reduces Agent Costs, Qwen3.8-Max-0902 Breaks Programming Benchmarks, Gemini Agentic Video Revolutionizes Long-Form Video Retrieval"},"content":{"rendered":"<h2 id=\"topic-634ecc425c\">Claude Fable 5.1 Brings Long-Task Capability, Cost, and Enterprise Privacy to the Forefront<\/h2>\n<p>Anthropic has released Claude Fable 5.1 and Mythos 5.1 for invited researchers; the core of this upgrade is not just benchmark scores, but enabling the model to autonomously execute complex work for longer durations. Official and user information shows that Fable 5.1 scored 52.6% on Terminal-Bench-Science 0.1, higher than Fable 5&#8217;s 24.7%; it scored 55.8% on Terminal-Bench 4.0, and Mythos 5.1 scored 60.9%, with the model having run autonomously for 38 consecutive hours in testing.<\/p>\n<p>Another change directly impacting practical use is the reduction of Cache Read pricing from $1 per million tokens to $0.25 per million tokens. Anthropic estimates total costs for typical workloads will drop by about 25%, and for highly agentic workloads by up to 45%. Concurrently, the Enterprise EFS plan keeps data within the customer&#8217;s own cloud and adds monitoring for agent risk patterns; however, blogger tests and discussions also point out that Fable 5.1&#8217;s subscription quota may deplete quickly, and API costs cannot be judged solely by the cache unit price.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@claudeai: <a href=\"https:\/\/x.com\/claudeai\/status\/2094848572143407483\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/claudeai\/status\/2094848572143407483<\/a><\/li>\n<li>@dotey: <a href=\"https:\/\/x.com\/dotey\/status\/2094854620732375335\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/dotey\/status\/2094854620732375335<\/a><\/li>\n<li>@alexalbert__: <a href=\"https:\/\/x.com\/alexalbert__\/status\/2094889286990446769\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/alexalbert__\/status\/2094889286990446769<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-eb40266419\">Qwen3.8-Max-0902 Refreshes the Agent Programming Competition with Continued Post-Training<\/h2>\n<p>Alibaba has released Qwen3.8-Max-0902; its foundation remains the 2.4T parameter, 1M context model, but it has undergone continued post-training focused on Coding and Cowork. The official pricing is $2 per million tokens for input, $6 for output, with explicit cache hits at $0.17 and implicit cache hits at $0.25; in Code Arena WebDev, this version temporarily ranks first with a score of 1691, surpassing the previous version by 22 points and Claude Opus 5 Max by 3 points.<\/p>\n<p>Comparisons by bloggers on the same series also show TerminalBench improving from 11.3 to 29.0, DeepSWE from 56.6 to 69.3, and QwenSWEBench V2 reaching 70.0. The value of this case lies in the fact that frontier models are no longer iterating solely by &#8220;major versions&#8221;; post-release data, reinforcement learning, and agent environment training are becoming the primary factors in continuously widening the performance gap.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@Alibaba_Qwen: <a href=\"https:\/\/x.com\/Alibaba_Qwen\/status\/2094968708288680276\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/Alibaba_Qwen\/status\/2094968708288680276<\/a><\/li>\n<li>@LufzzLiz: <a href=\"https:\/\/x.com\/LufzzLiz\/status\/2095085997214384624\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/LufzzLiz\/status\/2095085997214384624<\/a><\/li>\n<li>@cellinlab: <a href=\"https:\/\/x.com\/cellinlab\/status\/2095007265086668957\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/cellinlab\/status\/2095007265086668957<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-850c1ee31c\">Gemini API&#8217;s Agentic Video Shifts Long-Video Retrieval from &#8220;Frame-by-Frame Scanning&#8221; to Adaptive Forensics<\/h2>\n<p>Google has introduced Agentic Video to the Gemini API: the model can autonomously decide when to fast-forward, slow down, or rewind based on the question, and choose to focus on listening to audio, reading subtitles, or examining the visuals, rather than sampling frames at fixed intervals from long videos. A relevant Google executive stated that this processing method can reduce token consumption by up to 88% while improving result quality.<\/p>\n<p>Visible applications include locating scene changes, searching for specific events in hours-long videos, performing high-frame-rate anomaly detection on key time windows, and tracking actions and object counts over time. It is currently primarily accessible via API, with plans for future integration into YouTube&#8217;s Ask YouTube feature; its practical significance for video editing, long-video Q&amp;A, and visual quality inspection is more defined than simply adding a &#8220;watch video&#8221; entry point.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@OfficialLoganK: <a href=\"https:\/\/x.com\/OfficialLoganK\/status\/2094843143623680264\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/OfficialLoganK\/status\/2094843143623680264<\/a><\/li>\n<li>@Gorden_Sun: <a href=\"https:\/\/x.com\/Gorden_Sun\/status\/2095011424020128059\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/Gorden_Sun\/status\/2095011424020128059<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-00df51c5a8\">World Labs Atlas Advances Video Generation to Controllable 3D World Reconstruction<\/h2>\n<p>World Labs has released Atlas, officially described as a multimodal world model trained from scratch. It places text, images, video, camera poses, and 3D depth into a unified spatial context. It can reconstruct a scene from one or multiple images and generate videos up to 1 minute long at 1440p resolution, along with novel viewpoints, according to specified positions, angles, and motion trajectories.<\/p>\n<p>Independent experiences and project materials indicate that Atlas can also convert a small number of smartphone photos into point clouds or 3D Gaussian Splats and recreate &#8220;bullet time&#8221; effects using dynamic videos shot with several phones. In the robotics direction, it can generate RGB and depth data for virtual environments for training and testing. It is currently in Early Access for a limited number of partners. Its primary value lies in integrating cinematic camera control, spatial reconstruction, and robot simulation into the same model framework.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@theworldlabs: <a href=\"https:\/\/x.com\/theworldlabs\/status\/2094839756329041984\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/theworldlabs\/status\/2094839756329041984<\/a><\/li>\n<li>@drfeifei: <a href=\"https:\/\/x.com\/drfeifei\/status\/2094840371675283673\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/drfeifei\/status\/2094840371675283673<\/a><\/li>\n<li>@Gorden_Sun: <a href=\"https:\/\/x.com\/Gorden_Sun\/status\/2095093766327795718\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/Gorden_Sun\/status\/2095093766327795718<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-81389b546d\">OpenAI Elevates Astra&#8217;s Cybersecurity Capability to Critical, Simultaneously Lists Monitorability as a Constraint<\/h2>\n<p>In its release announcement, OpenAI confirmed that Astra has reached the Critical threshold for cybersecurity capability according to the Preparedness Framework, and stated that some development was delayed to strengthen isolation, network permissions, and monitoring. Sam Altman&#8217;s public statement also emphasized that Astra&#8217;s training is complete, but subsequent models will only advance when safety and alignment work is sufficient.<\/p>\n<p>Visible discussions surrounding Astra focus on two tensions: on one hand, reports indicate it can use deeper internal loops to increase reasoning computation, but this may leave more thinking in hidden states; on the other hand, OpenAI&#8217;s Chief Scientist responded that the computational graph depth of current frontier models remains within roughly twice that of GPT-4, and explicitly acknowledged that Chain-of-Thought monitorability is fragile and deteriorating. It is confirmed that safety capability and monitoring capability have simultaneously become release thresholds; this upgrade cannot be evaluated by a single benchmark alone.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@OpenAI: <a href=\"https:\/\/x.com\/OpenAI\/status\/2094885578173260259\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/OpenAI\/status\/2094885578173260259<\/a><\/li>\n<li>@sama: <a href=\"https:\/\/x.com\/sama\/status\/2094934592062959832\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/sama\/status\/2094934592062959832<\/a><\/li>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2095048557430727063\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2095048557430727063<\/a><\/li>\n<li>@derrickcchoi: <a href=\"https:\/\/x.com\/derrickcchoi\/status\/2094954736470614151\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/derrickcchoi\/status\/2094954736470614151<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-e920e74772\">Uber&#8217;s Software Factory Case Shows Agents Moving from Personal Assistants into Full R&amp;D Pipelines<\/h2>\n<p>A summary by a blogger of a sharing session from Uber&#8217;s engineering team shows that from February to mid-August 2026, the weekly active users of Uber&#8217;s internal Agent product grew 7-fold, and weekly request volume grew 9.4-fold, while the cost per session decreased by 52% from its peak; over 70% of Pull Requests were attributed to local or cloud-based Agents, with over 30,000 development-related Skills executed daily.<\/p>\n<p>This system covers code review, CI self-healing, end-to-end visual validation, alert triage, bug debugging and maintenance. It is managed by a unified model gateway handling identity, privacy, budget, and auditing, and connects processes via an MCP gateway, isolated development environments, a Skill marketplace, and a Context Graph. The key lesson it provides is not &#8220;giving employees a chatbot,&#8221; but embedding Agents into auditable, reusable production systems with human oversight.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@xiaohu: <a href=\"https:\/\/x.com\/xiaohu\/status\/2095148956460396751\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/xiaohu\/status\/2095148956460396751<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-e87e413663\">WorkBuddy Open Platform Begins Competing for the Identity and Permission Layer of Cross-Device Agents<\/h2>\n<p>Tencent launched the WorkBuddy open platform on September 2nd. Information from the first batch, as relayed by bloggers, includes over 100 partners, 9 co-branded hardware products, and over 30 industry applications, with open access for Buddy apps, Experts, Skills, Connectors, and hardware integration. Devices like glasses, recording cards, headphones, and keyboards\/mice are responsible for data collection and triggering, while WorkBuddy handles understanding, planning, tool invocation, and execution.<\/p>\n<p>Published interface documentation also shows that authorized third-party applications can create and read cloud tasks, and can also send messages to the local computer assistant and read its online status and message history. Therefore, what the platform is competing for is not just an office entry point, but rather identity, permissions, memory, task status, and distribution. Subsequent observation is still needed regarding cross-device task success rates, partner retention, developer revenue, and permission incident rates. The scale of the ecosystem cannot be confirmed by the initial number of integrations alone.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@LufzzLiz: <a href=\"https:\/\/x.com\/LufzzLiz\/status\/2095061433851797960\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/LufzzLiz\/status\/2095061433851797960<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-7673737df5\">China&#8217;s AI Governance Continues Extending Upstream from Content Handling to Models and Agent Applications<\/h2>\n<p>A blogger relayed progress from the second phase of the &#8220;Clear and Bright \u00b7 Rectifying AI Application Misconduct&#8221; campaign announced by the Cyberspace Administration of China (CAC), stating that the actions have cumulatively cleaned up over 5.61 million pieces of illegal and non-compliant information, investigated and dealt with over 49,000 accounts, and handled over 2,400 non-compliant websites and applications. The governance targets are seen to cover AI-altered and fake disaster information, impersonation via face\/voice swapping, vulgar and violent content, content violating regulations concerning minors, as well as AI-managed accounts and online trolls.<\/p>\n<p>The same information also mentioned that platforms like Doubao, Yuanbao, Qianwen, and Wenxin Yiyan have been required to strengthen original training data review, restrict non-compliant outputs, and label generated content; social platforms are upgrading multimodal content review, facial and voiceprint databases, while phone manufacturers and app stores are strengthening random inspections and reviews of AI apps. According to this public relay, the governance chain has expanded from &#8220;deletion after posting&#8221; to training data, model outputs, Agent behavior, and application distribution, but the specific implementation effectiveness should still be based on official subsequent disclosures.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2095001785266291083\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2095001785266291083<\/a><\/li>\n<\/ul>\n<p>Stats: Scanned timeline count=600 Number of bloggers matched=66 Total matched tweets=405 Weighted tweet score=297.75 Original tweet count=150 RT tweet count=108 Crawl attempt count=4 Boundary coverage status=tail_confidently_crossed_target_boundary<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Today&#8217;s AI field focuses on the simultaneous optimization of model capabilities and costs, as well as breakthrough progress in video understanding and generation technologies.<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[19],"class_list":["post-1678","post","type-post","status-publish","format-standard","hentry","category-brief","tag-x--ai-"],"_links":{"self":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts\/1678","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/comments?post=1678"}],"version-history":[{"count":0,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts\/1678\/revisions"}],"wp:attachment":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/media?parent=1678"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/categories?post=1678"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/tags?post=1678"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}