{"id":1551,"date":"2026-08-06T09:04:15","date_gmt":"2026-08-06T01:04:15","guid":{"rendered":"https:\/\/blog.liu-qi.cn\/2026\/08\/06\/x-daily-2026-08-05\/"},"modified":"2026-08-06T09:04:15","modified_gmt":"2026-08-06T01:04:15","slug":"x-daily-2026-08-05","status":"publish","type":"post","link":"https:\/\/en.blog.liu-qi.cn\/2026\/08\/06\/x-daily-2026-08-05\/","title":{"rendered":"X Platform August 5 AI Brief | AI Model Benchmarks Become Key for Selection, Agent Infrastructure and Design Collaboration Advance"},"content":{"rendered":"<h2 id=\"topic-b1822bb3de\">Benchmarking Seedance 2.5 vs. MiniMax H3: Cost and Quality Comparisons Now Possible with Identical Prompts<\/h2>\n<p>Practical tests using the same set of prompts have quantified the differences between Seedance 2.5, MiniMax H3, and happyhorse 1.1 across measurable metrics like resolution, price, performance, and continuity. Tool selection is shifting from &#8220;watching demos&#8221; to &#8220;task-based evaluation.&#8221; @joshesye compared H3 and Seedance 2.5 using identical video prompts: H3 outputs 2K, Seedance 2.5 outputs 720P, with recorded usage of 180 credits vs. 390 credits, noting H3&#8217;s cost could be as low as \u00a50.53 per second. @LufzzLiz&#8217;s first-round tests of Seedance 2.0, happyhorse 1.1, and MiniMax H3 noted differences in acting, night scene atmosphere, and key details, with all three potentially prioritizing major plot points over nuanced performance and continuity.<\/p>\n<p>Such practical tests are more useful for selection: first fix the prompt and reference image, then separately record character consistency, shot continuity, resolution, generation cost, and failure samples. Models should not be judged solely by the feel of a single finished clip.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@joshesye: <a href=\"https:\/\/x.com\/joshesye\/status\/2084852295607484730\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/joshesye\/status\/2084852295607484730<\/a><\/li>\n<li>@LufzzLiz: <a href=\"https:\/\/x.com\/LufzzLiz\/status\/2084818491430076693\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/LufzzLiz\/status\/2084818491430076693<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-060e33a807\">Cloudflare Turns Agent &#8220;Money&#8221; and Email Operations into Callable Infrastructure<\/h2>\n<p>Cloudflare officially introduced Wallets for storing stablecoins, purchasing services, and receiving funds; bloggers further explain this as configuring small budgets for Agents to autonomously trial APIs. @xiaohu gave an example: funding an Agent with about $10 can handle many API trial and payment hurdles. @LufzzLiz simultaneously cautioned that real fund custody, security audits, refund mechanisms, fees, and regional policies are not yet mature, making it more suitable for reserving names or testing on testnets for now.<\/p>\n<p>Separately, an independent developer released Mailworker based on Cloudflare Email and Workers, aiming to provide a multi-product, multi-domain inbox for Agents to read emails, draft replies, with human approval for sending via a dashboard. The project is still in alpha, and Cloudflare Email is also labeled beta, with the author explicitly recommending testing on domains not bound to a primary MX record and not deploying directly to production.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@Cloudflare: <a href=\"https:\/\/x.com\/Cloudflare\/status\/2084648084131242402\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/Cloudflare\/status\/2084648084131242402<\/a><\/li>\n<li>@xiaohu: <a href=\"https:\/\/x.com\/xiaohu\/status\/2084839042638872866\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/xiaohu\/status\/2084839042638872866<\/a><\/li>\n<li>@LufzzLiz: <a href=\"https:\/\/x.com\/LufzzLiz\/status\/2084676842146152790\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/LufzzLiz\/status\/2084676842146152790<\/a><\/li>\n<li>@iguangzhengli: <a href=\"https:\/\/x.com\/iguangzhengli\/status\/2084953958154682693\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/iguangzhengli\/status\/2084953958154682693<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-f9b333565e\">Open Design Enters Codex Official Plugin Directory, Agents Begin Collaborating on Real-Time Canvases<\/h2>\n<p>The official Open Design account announced its Codex plugin&#8217;s entry into the official plugin directory, offering a real-time design canvas directly within Codex. @tuturetom added that the plugin can interoperate with the Codex ecosystem while allowing model switching within Open Design with maintained context continuity. The project also previewed a full demo, an offline event in Hong Kong on August 16, and arrangements for free, unrestricted use of DeepSeek V4 Flash.<\/p>\n<p>The value of this information lies in the workspace interface expanding from pure text code editing to visual artifact collaboration. The Open Design community has also seen continuously updated books on local-first, Agent-native design workflows, indicating that this toolset is beginning to accumulate tutorials and usage methods, though the plugin&#8217;s actual effectiveness still requires testing with specific tasks.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@OpenDesignHQ: <a href=\"https:\/\/x.com\/OpenDesignHQ\/status\/2084984244762210484\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/OpenDesignHQ\/status\/2084984244762210484<\/a><\/li>\n<li>@tuturetom: <a href=\"https:\/\/x.com\/tuturetom\/status\/2084986587088110004\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/tuturetom\/status\/2084986587088110004<\/a><\/li>\n<li>@tuturetom: <a href=\"https:\/\/x.com\/tuturetom\/status\/2084938373035065652\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/tuturetom\/status\/2084938373035065652<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-31201ba258\">Qwen 3.8-Max Brings Capability Release, Open-Source Preview, and Price Competition Simultaneously<\/h2>\n<p>Alibaba Cloud officially released Qwen 3.8-Max, calling it the most capable model in the Qwen family to date, with a scale of 2.4 trillion parameters, covering programming, work, research, and long tasks. The same announcement previewed opening the Qwen 3.8-Max weights next week and open-sourcing the Qwen 3.8-27B weights. @NousResearch stated that Qwen 3.8-Max is already integrated with Hermes Agent and offers a 20% discount.<\/p>\n<p>Significant promotions also appeared on the pricing side: NousResearch&#8217;s Portal offers 20% off most models, 50% off GPT-5.6 Terra and Luna, and 90% off DeepSeek V4 Flash 0731. What can be confirmed here are the prices and integration information released by the accounts, not model quality rankings or long-term pricing strategies.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@alibaba_cloud: <a href=\"https:\/\/x.com\/alibaba_cloud\/status\/2084926666774905245\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/alibaba_cloud\/status\/2084926666774905245<\/a><\/li>\n<li>@NousResearch: <a href=\"https:\/\/x.com\/NousResearch\/status\/2084680562300862514\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/NousResearch\/status\/2084680562300862514<\/a><\/li>\n<li>@NousResearch: <a href=\"https:\/\/x.com\/NousResearch\/status\/2084680563756286220\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/NousResearch\/status\/2084680563756286220<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-c7caa2a665\">News on DeepSeek, Kimi, and SSI Funding and New Models Should Still Distinguish Between Reports and Confirmations<\/h2>\n<p>@MaxForAI relayed a report from &#8220;Caijing&#8221; stating that DeepSeek&#8217;s second funding round, paused due to a meeting recording leak incident, has restarted, planning to raise \u00a550 billion with a pre-money valuation of about \u00a5500 billion. The same account also stated that Moonshot AI is advancing a Series G funding round with a valuation of about $50 billion. The post interprets this as capital shifting from &#8220;can it be used&#8221; to betting on whether domestic models can replace overseas closed-source models, but these figures in this batch of visible content are media reports relayed by bloggers, not official confirmations from the relevant companies.<\/p>\n<p>Regarding SSI, @MaxForAI, based on an investor podcast, relayed that it may release its first model in August, mentioning connections to continual learning, sample-efficient learning, and NVIDIA collaboration. The post explicitly writes that SSI has not announced model capabilities, architecture, or a release date, and &#8220;has cracked continual learning&#8221; is merely external speculation. Trackable facts are the podcast claims and related collaboration clues, not that the model has been released or the technical route has been validated.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2084907288918438326\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2084907288918438326<\/a><\/li>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2084905969394532536\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2084905969394532536<\/a><\/li>\n<li>@MaxForAI: <a href=\"https:\/\/x.com\/MaxForAI\/status\/2084906051909140624\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MaxForAI\/status\/2084906051909140624<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-ee591778bc\">Skills Shift from &#8220;The More, the Better&#8221; to Manual Curation, Testing, and Monetizable Knowledge<\/h2>\n<p>@oran_ge stated that their team encountered degraded Agent performance after a user installed ten thousand Skills, leading to collaboration with @op7418 to establish a curated site emphasizing manual selection, testing, and recorded tutorials. @op7418&#8217;s introduced CoLa Skills Store also provides sources, explanations, and Chinese localization, matching Skills to user needs to avoid loading the entire massive library into an Agent at once. The key signal here is that Skill value is beginning to be determined by usability, safe handling, explanation cost, and maintenance quality.<\/p>\n<p>Knowledge products are also monetizing along the same direction. @MANISH1027512 announced the launch of the &#8220;Academy&#8221; on the VSC community, with the first course being a GPT-IMAGE 2 tutorial from long-term practical experience and manual writing, later stating that course sales approached \u00a520,000 within 12 hours of listing. This sales figure is self-reported by the publisher, indicating a case within the community of paying for verified methods, but cannot be extrapolated to the entire market size.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@oran_ge: <a href=\"https:\/\/x.com\/oran_ge\/status\/2084933132416176201\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/oran_ge\/status\/2084933132416176201<\/a><\/li>\n<li>@op7418: <a href=\"https:\/\/x.com\/op7418\/status\/2084929983924023560\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/op7418\/status\/2084929983924023560<\/a><\/li>\n<li>@MANISH1027512: <a href=\"https:\/\/x.com\/MANISH1027512\/status\/2084926617349210607\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MANISH1027512\/status\/2084926617349210607<\/a><\/li>\n<li>@MANISH1027512: <a href=\"https:\/\/x.com\/MANISH1027512\/status\/2084994221799587981\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/MANISH1027512\/status\/2084994221799587981<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-40a3295baf\">Model Engineering Begins Discussing Structured Input, Context Management, and Data Quality on the Same Level<\/h2>\n<p>@Gorden_Sun, based on a ByteDance Seed team paper, explained that simply lengthening natural language prompts mainly increases redundancy, while using JSON structures to express information like position, depth, pose, material, and background is more likely to improve text-to-image quality. @dongxi_nlp&#8217;s compiled research clues point to three separate issues: random rewards may amplify existing preferences without generating real capability, incorrect labels reinforce errors and weaken exploration, Agents need autonomous management of short-term context and long-term external memory, and pre-training data can be processed based on sample-level decisions.<\/p>\n<p>This content remains bloggers&#8217; summaries of papers and projects and cannot directly substitute for the papers&#8217; experimental conclusions. However, they collectively point to a specific change: bottlenecks for Agents and generative models lie not only in model size but also in prompt structure, context lifecycle, reward signals, and data filtering.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@Gorden_Sun: <a href=\"https:\/\/x.com\/Gorden_Sun\/status\/2084950115089834022\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/Gorden_Sun\/status\/2084950115089834022<\/a><\/li>\n<li>@dongxi_nlp: <a href=\"https:\/\/x.com\/dongxi_nlp\/status\/2084768262534173033\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/dongxi_nlp\/status\/2084768262534173033<\/a><\/li>\n<li>@dongxi_nlp: <a href=\"https:\/\/x.com\/dongxi_nlp\/status\/2085016112639377415\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/dongxi_nlp\/status\/2085016112639377415<\/a><\/li>\n<li>@dongxi_nlp: <a href=\"https:\/\/x.com\/dongxi_nlp\/status\/2085015920544534682\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/dongxi_nlp\/status\/2085015920544534682<\/a><\/li>\n<\/ul>\n<h2 id=\"topic-4471e54da8\">Agent Applications Show More Specific Internal Use and Real-Environment Test Cases<\/h2>\n<p>@xiaohu introduced Stripe&#8217;s internal Agent practice: an engineer spent a week creating an AI assistant named Kai for the entire company, aiming to allow non-engineering employees to use an &#8220;AI colleague&#8221; familiar with internal operations without extra configuration. This is a blogger&#8217;s summary of an article, supporting the case judgment that &#8220;internal knowledge and processes are encapsulated into Agents,&#8221; but cannot confirm the specific company-wide usage ratio.<\/p>\n<p>@steipete explained that to automate testing of OpenClaw&#8217;s iMessage integration, they connected Codex to a remote KVM with video capability because iMessage is unreliable in virtual machines, and features like read receipts involve disabling SIP. The former showcases the target users for Agents within organizations, the latter showcases end-to-end testing constraints for Agents on real systems. The gap mainly stems from environment, permissions, and device state, not simply model capability.<\/p>\n<p>Sources:<\/p>\n<ul>\n<li>@xiaohu: <a href=\"https:\/\/x.com\/xiaohu\/status\/2085013960991105112\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/xiaohu\/status\/2085013960991105112<\/a><\/li>\n<li>@steipete: <a href=\"https:\/\/x.com\/steipete\/status\/2084988316324397312\" target=\"_blank\" rel=\"noopener noreferrer\">https:\/\/x.com\/steipete\/status\/2084988316324397312<\/a><\/li>\n<\/ul>\n<p>Stats: Timeline posts scanned=360 Bloggers matched=34 Total tweets matched=205 Weighted tweet score=165.25 Original tweets=93 Retweeted tweets=36 Fetch attempts=2 Boundary coverage status=tail_confidently_crossed_target_boundary<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Today&#8217;s focus includes cost and quality benchmarks for AI video models, Cloudflare&#8217;s infrastructure for Agent funding and email management, and real-time collaboration integration for design tools and code editors.<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[5],"tags":[19],"class_list":["post-1551","post","type-post","status-publish","format-standard","hentry","category-brief","tag-x--ai-"],"_links":{"self":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts\/1551","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/comments?post=1551"}],"version-history":[{"count":0,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/posts\/1551\/revisions"}],"wp:attachment":[{"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/media?parent=1551"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/categories?post=1551"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/en.blog.liu-qi.cn\/index.php\/wp-json\/wp\/v2\/tags?post=1551"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}