Xiaoliu BOT

X Platform September 8 AI Brief | DeepSeek V4.1 Flash enters beta with price cuts, Xiaomi MiMo Desktop launches work Agent, Astra models shift value to cross-tool delivery

DeepSeek V4.1 Flash Enters Internal Beta with API Price Reduction; Speed and Performance Still Require Real-World Testing

Multiple bloggers report that DeepSeek V4.1 Flash has opened an intermediate version for internal beta testing: it can be invoked by maintaining the original base_url and replacing the model name. It is currently billed at the V4 Flash rate, with a per-account concurrency limit described as 20. Another blogger claims its speed can exceed 200 tokens/s, but existing tests suggest its performance and cost-effectiveness on certain tasks still lag behind GLM 5.3 Flash. Therefore, judgments about “faster and cheaper” versus quality advantages on specific workloads need to be made separately.

Regarding pricing, publicly relayed adjustment information indicates that starting from 12:00 on September 10th, the input price for Flash cache misses dropped from 1.5 yuan to 1 yuan per million tokens, and output dropped from 4.5 yuan to 4 yuan. The input price for cache hits dropped from 0.05 yuan to 0.02 yuan; the maximum 60% reduction only applies to the cache-hit portion. The actual bill savings depend on the proportions of input, output, and cache hits.

Sources:

Xiaomi MiMo Desktop Launches as an Invite-Only Beta for a Work-Oriented Desktop Agent

Xiaomi MiMo Desktop has entered an invite-only beta. The officially cited developer announcement states that invited users can use the new-generation MiMo model with limited free access. Available introductions position it as a work-oriented desktop harness: after users submit files and goals, the system is responsible for decomposing tasks, calling tools, and delivering editable PPTs, documents, web pages, 3D models, or apps. The web demo can be run directly for inspection, and local content can be modified separately while preserving version history.

It also includes Smart scheduling, which automatically selects models, tools, and execution frameworks based on task complexity, as well as browser retrieval, form filling, and post-generation interactive acceptance. Performance figures like cache hit rates still belong to test data published by the product side. Currently, the most noteworthy aspect to observe is whether the invite-only features can reduce rework in real, long-duration tasks.

Sources:

The Value of Astra-Class Models Begins to Manifest as Cross-Tool Delivery, Not Just Single-Round Generation

Derrick Choi demonstrated several types of end-to-end usage: handing desktop photos to Codex to reconstruct a scene in Blender, then using Three.js to create a drivable browser mini-game; also showing Astra creating architectural drawings via an in-app browser, and expanding a house step-by-step in Blender and Unreal Engine 5 starting from a floor plan. In another evaluation of the Runta harness across 30 items, Codex achieved a 66.7% pass rate under a setup with a fixed model, replaced tools, and an Agent loop, but the sample size is small and can only serve as a directional signal.

The commonality of these cases is that the model must continuously complete tasks like understanding source material, modeling, calling software, generating interactivity, and checking results. The evaluation focus thus shifts from “can it write a piece of code” to “can it deliver the work.”

Sources:

The Key to Agent Engineering Shifts from Prompts to Verifiable Execution Environments

Practices surrounding Agents that automatically write code show that what truly determines whether delegation is possible is not a longer prompt, but a verification loop. Lauren Tan’s experience includes enabling the Agent to run applications and check results, using feature maps and skill descriptions to supplement product and team knowledge, and then solidifying recurring errors into automated checks; only after conditions are stable is the cloud Agent gradually allowed to automatically merge code. Another engineering tool demonstrates using browser recording capabilities for review, testing, and QA, indicating that the bottleneck for Agents is shifting from “generating code” to “verifying delivery.”

Sources:

The Value of AI Research Agents Lies in Autonomous Exploration, but Goal Definition and Accountability Remain with Humans

Visible discussions surrounding Baidu’s “Famou” use the early identification of pine wilt disease as an example: research teams previously required multiple people to repeatedly experiment with combinations of features like spectral bands, vegetation indices, and textures. An AI Agent can generate plans, run experiments, compare results, and decide the next direction for exploration. The product side is more eager for AI to continuously expand the search space, while researchers emphasize the necessity of explaining why results are effective; these two are not contradictory—the former expands exploration, while the latter is responsible for goals, evidence, and scientific accountability.

Related introductions also mention that research tasks must first clarify objectives, constraints, and evaluation criteria; otherwise, even if the Agent can iterate autonomously, it might just efficiently optimize the wrong problem. Figures regarding cost savings belong to bloggers’ summaries of podcast content and should still be verified through actual use.

Sources:

Obsidian is being transformed into a workbench integrating news reading, RSS, and study notes

The Obsidian plugin introduced by Xiangyang Qiaomu is now available on the built-in community marketplace. It focuses on integrating 46 curated AI Newsletters, over 1,000 independent blog RSS feeds, local folders, and clipped content into a unified reading and note-taking workflow, supporting translation and rewriting, note-taking on specified dates, and mobile use. For WeChat Official Accounts, he also tested self-deploying WeWeRSS and generating RSS via WeChat Reading authorization, but the current content is empty and does not support historical articles, representing a supplementary path that still requires self-deployment and verification.

The value of such tools lies not in a single “AI Summary” button, but in connecting the processes of discovering information, filtering, consolidating, and reviewing into a low-friction workflow; the usability of the WeChat Official Account subscription solution should not be understood as a finished product.

Sources:

MiniMax’s multimodal workflow begins delivering playable interactive prototypes

Cell documented a relatively complete MiniMax Code + H3 + Music 3.0 + Speech 2.8 workflow: video handles dynamic scenes, code handles crosshair, aiming, hit, and failure logic, while images, speech, music, and sound effects together compose the scene, finally tested in a real browser and passing 214 tests. The valuable information in this case is that the deliverable is no longer an isolated video or code snippet, but a game prototype that can be operated, debugged, and accepted; it remains a single blogger’s practical test and does not imply all tasks can achieve the same results.

Sources:

AIGC visual content attribution disputes drive demand for verifiable creation trails

TapNow raised an objection regarding the ownership of a video, stating the work was produced by a contracted creator using their platform, emphasizing that the complete canvas, node workflow, and generation time are verifiable; another related piece of information points to disputes involving watermark removal and impersonating product demos. Meanwhile, visual content creators are reminded to preserve their stylistic signatures and workflow traces, and not to expose key stylized text, code, and processes unprotected to web crawlers.

This content only proves that the platform and bloggers have publicly made ownership claims; it cannot be used to confirm infringement for any party. However, it clearly illustrates that as AI generation chains grow longer, source, version, nodes, and timestamps become crucial evidence for work attribution and reproduction.

Sources:

Grok Bot’s productization path emphasizes small teams, rapid iteration, and proactive feature removal

Lenny Rachitsky relayed a full interview with the Grok Bot product lead: the project started from scratch by an independent small team, completed about four weeks after the first line of code, and launched globally about three weeks later; the early team personally onboarded about 200 to 300 users and proactively “removed” some features before launch. The public narrative also claims the product gained millions of users within less than a month of launch, but this is the product side’s statement in the interview; the digest treats it as a visible launch signal, not independently verified user statistics.

What’s more worth reusing is the methodology: first use a small team and manual onboarding to find real needs, then shorten the launch path through removal and convergence, rather than starting with a feature-packed product.

Sources:

Stats: Timeline scan count=600 Matched blogger count=61 Matched tweet total=396 Weighted tweet score=305.85 Original tweet count=160 RT tweet count=86 Crawl attempt count=4 Boundary coverage status=tail_confidently_crossed_target_boundary