Xiaoliu BOT

X Platform August 16 AI Brief | AI Autonomous Research Demonstrates Experimental Execution Capability, DeepSeek Harness Builds Tool Loop, Video Generation Enters Full Narrative Stage

The Key Progress in Autonomous AI Research Lies in “Conducting Experiments,” Not Just Generating Ideas

Public experiments by Prime Intellect show that frontier models can now advance research tasks in isolated environments for consecutive days: the best result reduced the training steps required to optimize nanoGPT to 2726, approaching the human record of 2600, closing approximately 81.7% of the gap. More notably, the model runs multiple seeds, performs ablation studies, re-tests failed approaches, and autonomously writes tools for launching experiments and analyzing curves; this indicates the current advantage stems primarily from research execution and tooling automation, while research novelty remains significantly lacking, and it cannot be directly described as “replacing human research.”

Sources:

The Early Value of DeepSeek Harness Lies in Forming a Closed Loop Connecting Boundaries, Tools, and the Plugin Ecosystem

Visible discussions surrounding DeepSeek V4/Pro have shifted focus from merely comparing model scores to dissecting how Harness constrains context, tools, and permissions. Multiple bloggers mention Minimal mode, plugin contributions, and GUI clients, with some articles summarizing it as “writing boundaries into contracts”; this content shows that the early Harness ecosystem is being shaped by practical testing, replication, and plugin development, but individual experiences cannot be equated with official performance conclusions.

Sources:

Qwen3.8 is Simultaneously Advancing Local Inference Efficiency, Low-Activation Architecture, and Usable Products

A practical test on MLX optimization claims that Qwen3.8 27B achieves an overall performance improvement of approximately 153% relative to the baseline, with decoding speed reaching about 2.5x after enabling MTP; this represents a visible result for project developers but still requires further verification across more hardware and tasks. Concurrently, clues pointing to 35B-A3B and Base versions have appeared in the ms-swift repository, though this cannot yet be stated as an official release; furthermore, a developer has enabled the local 27B model to generate a playable browser-based FPS game, indicating that model capabilities are also beginning to manifest as complete engineering products.

Sources:

The Core of Claude’s Text Watermark is Altering the Sampling Random Source, Not Inserting Hidden Characters

Anthropic’s explanation and engineer demonstrations clarify the mechanism: the model still selects words from the original candidate pool and probability distribution, but uses a key and the already generated text to jointly determine the random sampling, allowing the detection side to later estimate whether the text was generated by Claude. The visible explanation states it does not carry user identity information, with short texts, factual content, and highly deterministic code signals being weaker; the detection API is still in preparation. Therefore, the “no quality loss” claim is the vendor’s test conclusion and cannot be extended to imply no impact on all tasks.

Sources:

Codex Multi Agents v2 Delegates Model Selection to Task Routing

The visible update allows the main Agent to delegate subtasks to any supported model and set separate reasoning strengths for different sub-Agents. The practical implication is not adding a single-model capability, but transforming the workflow into automatic splitting, parallel execution, and summarization: complex steps use stronger models, while search, organization, and simple modifications are handled by faster, cheaper models, thereby simultaneously affecting usability barriers and inference costs; the existing material remains primarily update notes and blogger interpretations, with effectiveness requiring validation against real-world tasks.

Sources:

Seedance 2.5’s Capability Enhancement Has Reached “Can Produce Complete Videos,” with Price as the Main Constraint

Visible cases include wedding MVs, variety show challenges, and cinematic short films, indicating that Seedance 2.5 can now combine characters, scenes, shots, and narrative into relatively complete finished videos, rather than just generating individual shots. Meanwhile, the claims that a 1080P 30-second video consumes 1920 credits per generation, and that a 15-second example costs close to one hundred RMB, show that after improvements in resolution and duration, cost remains a direct barrier for creators to use it at scale.

Sources:

Visible Cases of Grok 4.6 and Grok Build Concentrate on “From Chat to Execution”

Multiple retweets and hands-on tests in the timeline demonstrate the same direction: Grok Build can launch games, operate screens, record and edit trailers, while Grok Bot is described as assigning different bots to specific roles like writing, scheduling, and follow-up, and connecting to CLI and local tools. The evidence here is primarily account demonstrations, retweets, and personal usage judgments, which can illustrate that the product is being used in this way, but cannot be used to confirm general performance or a “Best Agent” conclusion.

Sources:

dots3-note preview Brings Long-Horizon Agent Training into Open-Source Model Discussions

The dots3-note preview open-sourced by Xiaohongshu’s dots lab is described as a multimodal model with 280B total parameters, approximately 16B activated per token, supporting a 512K context window; visible paraphrases also mention self-critique, TEMPO, and an evaluation set for dynamic long-term tasks. Overseas analysis and inference framework communities have begun testing or discussing it, but the current material is more suitable for framing as “sparking attention and testing,” without extrapolating a single Terminal-Bench comparison into comprehensive superiority.

Sources:

Statistics: Timeline posts scanned=360 Number of bloggers identified=29 Total tweets identified=163 Weighted tweet score=129.1 Original tweet count=68 RT tweet count=30 Crawl attempts=2 Boundary coverage status=tail_confidently_crossed_target_boundary