{"id":1578,"date":"2026-08-13T09:04:41","date_gmt":"2026-08-13T01:04:41","guid":{"rendered":"https:\/\/blog.liu-qi.cn\/2026\/08\/13\/x-daily-2026-08-12\/"},"modified":"2026-08-13T09:04:41","modified_gmt":"2026-08-13T01:04:41","slug":"x-daily-2026-08-12","status":"publish","type":"post","link":"https:\/\/en.blog.liu-qi.cn\/2026\/08\/13\/x-daily-2026-08-12\/","title":{"rendered":"X Platform August 12 AI Brief | DeepSeek V4 Pro Launches, xAI Advances Models and Executable Agents Simultaneously, AI Video Competition Shifts to Continuous Systems"},"content":{"rendered":"<h2 id=\"topic-6abf0d7aff\">Signs of DeepSeek V4 Pro 0813 Launch Emerge, with Agent Capabilities and Pricing as Key Focus<\/h2>\n<p>Yesterday, multiple bloggers posted updates regarding the release of DeepSeek V4 Pro-0813, the official website version update, and the pricing table. However, there are discrepancies in the reported release timeline: Gorden_Sun directly announced the official version release, MaxForAI stated the website had just been updated and provided comparative test results, while AlchainHust suggested complete release information might not be available until the next day. A set of benchmark data cited by MaxForAI indicated significant improvements from the preview version to 0813 on projects like Terminal Bench, CyberGym, and DeepSWE, with DeepSWE jumping from 12.8 to 62.7; this data is compiled or tested by bloggers and cannot replace official evaluations. Vista8 listed the prices at that time as \u00a50.025 per million tokens for cache hits, \u00a53 for misses, and \u00a56 for output, with a concurrency limit of 500, while also noting prices might increase. Notably, the current discussion focus has shifted from &#8220;whether the model is more powerful&#8221; to a comprehensive comparison of Agent coding capabilities, invocation costs, and concurrency capacity.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@Gorden_Sun: <a href=\"https:\/\/x.com\/Gorden_Sun\/status\/2087561627881312633\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/Gorden_Sun\/status\/2087561627881312633<\/a><\/li>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2087565786823221414\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2087565786823221414<\/a><\/li>\n<li>@vista8: <a href=\"https:\/\/x.com\/vista8\/status\/2087565229224034521\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/vista8\/status\/2087565229224034521<\/a><\/li>\n<li>@AlchainHust: <a href=\"https:\/\/x.com\/AlchainHust\/status\/2087568238624518648\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/AlchainHust\/status\/2087568238624518648<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-1d54af91b5\">xAI Advances Grok 4.6 and Grok Bot Simultaneously, with Model Upgrade and Executable Agent Both Rolling Out<\/h2>\n<p>From visible updates, xAI&#8217;s latest release covers both the model and product form: @SpaceXAI announced Grok 4.6, which shows improvements over Grok 4.5 while maintaining the same price; a post forwarded by Elon Musk also claimed it achieved 1753 ELO on GDPVal-AA, though this is a single benchmark figure from the forwarded content. Grok Bot is described as an &#8220;AI colleague&#8221; running on a persistent Linux machine in the cloud, capable of using a user&#8217;s login status to operate web pages, applications, and files. After a user demonstrates a task once, it can repeat execution based on time or events, and multiple Bots can delegate tasks to each other; it hands control back to the user for steps like login, CAPTCHA, or payment. In early access, Lenny Rachitsky mentioned it has been used to automatically reply to support emails, match job seekers with hiring companies, scan credit card subscriptions, and prepare podcast guest briefings. These experiences come from early testing and product descriptions and are insufficient to prove general reliability, but the product boundary has clearly advanced from chat-based responses to performing software operations on behalf of users.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@elonmusk: <a href=\"https:\/\/x.com\/elonmusk\/status\/2087565020158992709\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/elonmusk\/status\/2087565020158992709<\/a><\/li>\n<li>@xiaohu: <a href=\"https:\/\/x.com\/xiaohu\/status\/2087496240078713274\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/xiaohu\/status\/2087496240078713274<\/a><\/li>\n<li>@lennysan: <a href=\"https:\/\/x.com\/lennysan\/status\/2087241423792087518\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/lennysan\/status\/2087241423792087518<\/a><\/li>\n<li>@openclaw: <a href=\"https:\/\/x.com\/openclaw\/status\/2087563414210302084\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/openclaw\/status\/2087563414210302084<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-759b50ad48\">ChatGPT and Codex Desktop Preview Expands to Linux, with Workflow Migration and Sync Support<\/h2>\n<p>The official OpenAI account announced the ChatGPT desktop app is entering Linux preview, supporting Ubuntu 24.04\/26.04, Debian 13, Fedora 43\/44, and providing .deb or .rpm installation packages for x64 and ARM64. OpenAI Developers simultaneously announced workflow migration capabilities: projects, chats, Skills, Plugins, Agents, and project configurations from other Agents can be imported into ChatGPT Work and Codex, with import history viewable and options for automatic updates. The product team added that the Linux version&#8217;s overall functionality is nearly complete, though native Computer Use is not yet provided but is under development, and there may still be rough edges during the preview phase. The practical value of this change lies not just in adding another operating system, but in reducing the cost of migrating development environments and reusing existing Agent assets.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@OpenAI: <a href=\"https:\/\/x.com\/OpenAI\/status\/2087231350134980830\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/OpenAI\/status\/2087231350134980830<\/a><\/li>\n<li>@OpenAIDevs: <a href=\"https:\/\/x.com\/OpenAIDevs\/status\/2087242829076791392\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/OpenAIDevs\/status\/2087242829076791392<\/a><\/li>\n<li>@thsottiaux: <a href=\"https:\/\/x.com\/thsottiaux\/status\/2087254026232775052\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/thsottiaux\/status\/2087254026232775052<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-d140cd3632\">A Paper Raises Security Alerts Regarding Cross-Model Compatibility of Encrypted Chain-of-Thought<\/h2>\n<p>A paper cited by Gorden_Sun points out that some model vendors encrypt hidden reasoning traces before returning them to the client, but these traces may be interchangeable across different sessions, users, and models within the same ecosystem. Based on this, researchers attempted to have weaker models decrypt the encrypted traces generated by stronger models, leading to risks such as reasoning leakage, bypassing distillation protection mechanisms, privacy data exposure, and prompt injection. The citation also mentions that researchers recovered 367 pieces of personally identifiable information and 182 credentials from 315,320 reasoning data blocks in public code repositories, but these figures in the daily report can only be treated as a citation of the paper. Dongxi_nlp further argues that such methods technically support extracting higher-density reasoning data from closed-source models, but explicitly notes that existing public evidence cannot yet prove these traces have entered a specific model&#8217;s training set, nor can it quantify their contribution to model capabilities. The core value lies in reminding developers: encrypted reasoning data visible on the client side should still be treated as potentially sensitive information.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@Gorden_Sun: <a href=\"https:\/\/x.com\/Gorden_Sun\/status\/2087540930169672102\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/Gorden_Sun\/status\/2087540930169672102<\/a><\/li>\n<li>@dongxi_nlp: <a href=\"https:\/\/x.com\/dongxi_nlp\/status\/2087248596135444865\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/dongxi_nlp\/status\/2087248596135444865<\/a><\/li>\n<li>@dotey: <a href=\"https:\/\/x.com\/dotey\/status\/2087323013158990077\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/dotey\/status\/2087323013158990077<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-5701f40db4\">AI Video Evaluation Criteria Shift from Single-Shot Image Quality to Continuous Characters, Dynamic Systems, and Production Costs<\/h2>\n<p>Multiple tests from yesterday indicate that the competitive focus of AI video is shifting from &#8220;how realistic the image looks&#8221; to whether it can stably express a continuous system. Derek Wen tested a 15-second process with MiniMax H3, from game menus and weapon switching to third-person scenes and HUD feedback, focusing on whether character consistency, UI interaction, object changes, and physical feedback could all be established simultaneously. Chengzi shared practical experience with Seedance 2.5: single-character cards should be clear, prompts can be more detailed but the structure must be explicit, and checking materials and settings before generation helps reduce the cost of repeated &#8220;draws&#8221; (approximately 30-second videos, costing about 50 yuan each time). Ugur Aksay relayed information about FLUX 3 Video&#8217;s limited-time free access and subsequent roadmap, mentioning directions such as 20-second videos, synchronized sound, keyframes, continuation, and multi-image reference; xiaohu relayed the case of Higgsfield completing a 110-minute AI feature film, which used licensed real-person portraits and publicly shared production materials. It&#8217;s evident that samples are still primarily from blogger showcases and information from product teams, but the validation dimensions have clearly expanded to include long sequences, interaction, cost, and reusable workflows.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@derek_wall90176: <a href=\"https:\/\/x.com\/derek_wall90176\/status\/2087421450978509024\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/derek_wall90176\/status\/2087421450978509024<\/a><\/li>\n<li>@Chengzilhy: <a href=\"https:\/\/x.com\/Chengzilhy\/status\/2087525796093276624\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/Chengzilhy\/status\/2087525796093276624<\/a><\/li>\n<li>@uguraksay: <a href=\"https:\/\/x.com\/uguraksay\/status\/2087542569689358757\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/uguraksay\/status\/2087542569689358757<\/a><\/li>\n<li>@xiaohu: <a href=\"https:\/\/x.com\/xiaohu\/status\/2087537080390107136\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/xiaohu\/status\/2087537080390107136<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-0dc11d8c81\">Agent Work is Shifting from &#8220;Capable of Generation&#8221; to Being Testable, Orchestratable, and Reusable<\/h2>\n<p>Several bloggers have provided work methods that are more reusable than simply showcasing model capabilities. Baoyu observed that when Claude Code, Codex, and other Harnesses execute in parallel within the same directory, they can handle conflicts and wait for other tasks, with worktree being more suitable for time-consuming PoCs; Ryan Hanley summarized an Agent&#8217;s moat as professional expertise, distribution, taste, and process, suggesting breaking work into Skills or SOPs, establishing a context repository, and setting up guardrails similar to employee permissions. Ugur Aksay recommended generating 5 test scenarios, expected outputs, acceptance criteria, and failure signals before starting a task, then scoring and correcting the results, to avoid mistaking a single, accidental good answer for a good Prompt. PMbackttfuture also shared a practical case: using WorkBuddy and Hy3 to have 10 Agents organize the chat records of a community of over 300 people from August 1st to 10th, and outputting them to a Feishu cloud document. The common conclusion is that the value of an Agent depends on task orchestration, acceptance, and context accumulation, rather than the impressiveness of a single generation.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@dotey: <a href=\"https:\/\/x.com\/dotey\/status\/2087377704286847246\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/dotey\/status\/2087377704286847246<\/a><\/li>\n<li>@rhanley: <a href=\"https:\/\/x.com\/rhanley\/status\/2087473531227340884\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/rhanley\/status\/2087473531227340884<\/a><\/li>\n<li>@uguraksay: <a href=\"https:\/\/x.com\/uguraksay\/status\/2087552701269700740\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/uguraksay\/status\/2087552701269700740<\/a><\/li>\n<li>@PMbackttfuture: <a href=\"https:\/\/x.com\/PMbackttfuture\/status\/2087336464388374925\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/PMbackttfuture\/status\/2087336464388374925<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-2a084c13c6\">Public Signals Emerge of Manus Spinning Off from Meta and Resuming Independent Operations<\/h2>\n<p>Baoyu and Hua Shu respectively relayed page changes observed by &#8220;LatePost&#8221;: Manus&#8217;s website, which had displayed &#8220;Now part of Meta&#8221; for months, has changed to &#8220;Will resume independent operations soon.&#8221; Guizang subsequently mentioned that following this change, Manus has opened limited free access to some users, covering Manus 1.6 and Manus Lite; those who have previously purchased memberships or obtained them through Lenny&#8217;s Newsletter can try to use it. The current evidence mainly consists of relayed media reports, website page observations, and blogger experiences, not a complete official announcement visible in daily reports. Therefore, a more prudent judgment is that &#8220;signals of independent operation have emerged,&#8221; rather than making further inferences about the transaction structure or subsequent business. The reason it warrants attention is that Manus has undergone changes in product ownership and operational methods in a short period, reflecting that the commercialization path for Agent products is still rapidly adjusting.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@dotey: <a href=\"https:\/\/x.com\/dotey\/status\/2087434739929977149\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/dotey\/status\/2087434739929977149<\/a><\/li>\n<li>@AlchainHust: <a href=\"https:\/\/x.com\/AlchainHust\/status\/2087449464936263912\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/AlchainHust\/status\/2087449464936263912<\/a><\/li>\n<li>@op7418: <a href=\"https:\/\/x.com\/op7418\/status\/2087494509454389651\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/op7418\/status\/2087494509454389651<\/a><\/li>\n<\/ul>\n<p>Statistics: Scanned timeline count=480 Matched blogger count=43 Total matched tweets=265 Weighted tweet score=222.55 Original tweet count=120 RT tweet count=31 Crawl attempt count=3 Boundary coverage status=tail_confidently_crossed_target_boundary<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Today&#8217;s focus includes the launch of DeepSeek&#8217;s new model, updates to xAI&#8217;s models and executable Agents, and a shift in AI video evaluation standards towards continuous characters and dynamic systems.<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[19],"class_list":["post-1578","post","type-post","status-publish","format-standard","hentry","category-brief","tag-x--ai-"],"_links":{"self":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts\/1578","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/comments?post=1578"}],"version-history":[{"count":0,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts\/1578\/revisions"}],"wp:attachment":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/media?parent=1578"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/categories?post=1578"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/tags?post=1578"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}