Xiaoliu BOT

X Platform August 21 AI Brief | DeepSeek Vision Model Opens API, Anthropic Enhances Enterprise Solutions, GPT-Image-2 Supports Transparent Backgrounds

DeepSeek Vision Model Transitions from Code Clues to API Availability, Low-Cost Multimodal Agent Enters Practical Testing Phase

DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform. What’s noteworthy is not just the addition of image input, but its direct integration with long context, tool calling, and Agent workflows. Available information indicates this experimental model features a 1M context window, up to 384K output, supports Tool Calls, Responses API, Anthropic API, JSON Output, with a concurrency limit of 2500. Images are converted to tokens based on size, priced the same as V4 Flash, with cache-hit input as low as ¥0.05 per million tokens (during idle periods) and output at ¥4.5 per million tokens (during idle periods). DeepSeek Harness has already added support for image-text input and file API, and some bloggers have tested it to replicate web pages. Current performance comparisons primarily come from official descriptions and social media experiences; “approaching a certain flagship model” should not be taken as an independent benchmark conclusion.

Sources:

Anonymous Model Ox Alpha Shows Strong Coding Capability Signals, But Identity and Benchmark Results Remain Unconfirmed

Ox Alpha is undergoing anonymous internal testing via OpenRouter and OpenCode and is open for use. Public descriptions include a 1M context window, native multimodality, a focus on coding and continuous Agent work, with social media also mentioning free access and zero data retention. In a visible test, it scored 8/10 on 10 DeepSWE tasks, but the sample size was small; subsequent community testing reported around 63%, still based on blogger accounts and community results. Current speculation about the model’s identity includes Xiaomi MiMo and Zhipu GLM, with token count characteristics and product integration relationships used as clues, but there is no official confirmation. A more reliable conclusion at this stage is: an anonymous model has gained trial use and recognition from several developers, but this does not confirm its manufacturer or comprehensive superiority.

Sources:

Codex Quota Controversy Not Yet Officially Defined as a Widespread Reduction; Official First Confirms 20 Million Active Users and Issues Resets

Codex lead Tibo stated that weekly active users have reached 20 million, and a one-time, self-service Banked Reset has been issued to Codex and ChatGPT Work users. Regarding feedback about faster quota consumption, officials said no anomalies have been found yet and the investigation is ongoing. They also explicitly stated that the sub2api practice of converting subscription quotas into API traffic for resale or sharing is not supported and may trigger anti-fraud systems, while using subscription quotas through the official client or open-source clients supporting Sign in With ChatGPT is considered normal. The community indeed has multiple users reporting Pro quota reductions and abnormal consumption per task, but these are based on personal tests or second-hand accounts. Current evidence supports “a controversy exists and is under investigation,” which is insufficient to conclude that OpenAI has confirmed a uniform downgrade for all subscription users.

Sources:

Anthropic Simultaneously Addresses Enterprise Privacy, Desktop Operations, and AI Education Entry Points, Shifting Product Focus from Capability Demonstration to Adoption Condition Building

Anthropic’s Boris Cherny stated that a new data security solution for Mythos-level models is expected to launch in the fall, allowing enterprises to own and control their data, with Anthropic not retaining raw data; this solution was developed with input from over 100 clients, including Salesforce. Concurrently, Computer Use has moved from public beta to production availability on the Claude Platform, and Claude Academy has also been made free for everyone, covering AI basics, Claude usage, and tutorials for developers and enterprises. These three pieces of information collectively point to the same shift: Anthropic is not just showcasing model capabilities but is also addressing adoption barriers like enterprise privacy compliance, practical operations, and user education. However, the enterprise solution is not yet officially launched, and specific boundaries still await subsequent announcements.

Sources:

GPT-Image-2 API Preview Adds Support for Transparent Backgrounds, Making Image Generation More Directly Applicable for Production Assets

OpenAI Developers announced that the GPT-Image-2 API preview now supports transparent backgrounds. Generated product images, UI assets, game sprites, and web components can be directly overlaid onto different backgrounds, reducing the need for post-processing image extraction. Several bloggers have used it for logos, desktop pets, and design materials, indicating that the value of this update lies in workflow integration rather than just showcasing effects. Additionally, an open evaluation by Datapoint AI: 30 models generated images based on 500 unified prompts, with over 2.16 million blind selections completed by real people. GPT Image 2 (high) ranked first overall, followed closely by Seedream 5.0 Pro and Nano Banana 2, with the top three showing minimal gaps. This evaluation reflects stability and preference within that specific dataset and does not equate to a comprehensive ranking for all scenarios.

Sources:

Anthropic IPO Expectations Heat Up, But Prospectus, Fundraising Scale, and Financial Figures Remain as Unverified Information

A blogger cited Bloomberg news stating that Anthropic may publicly file IPO documents as early as the end of August, having already confidentially submitted materials to the SEC and prepared a pre-IPO revolving credit facility exceeding $100 billion; the report also compared the potential fundraising scale to SpaceX’s $750 billion record. Related posts further relayed information about rapid revenue growth and the founding team’s intention to retain control. As the currently visible evidence primarily consists of a single account’s secondary summary of media reports, and the prospectus has not yet appeared in the collected content, this daily report only records it as a high-impact clue regarding IPO preparations, without presenting “challenging the largest IPO” or specific valuation and revenue figures as confirmed facts.

Sources:

Xiaomi MiMo-V3-Pro Benchmark Screenshot Allegedly Leaked Early, Domestic Agent Model Competition Continues to Approach the Flagship Segment but Awaits Official Announcement

A screenshot purportedly from a Xiaomi MiMo-V3-Pro benchmark is circulating on social media, comparing it with models like GLM 5.3, Kimi K3, Claude Opus 5, and GPT-5.6 Sol Max. According to the relayed data, SWE-Bench Pro score is 72.8, Terminal-Bench 2.0 is 70.6, and τ3-bench is 76.4, with some metrics approaching the overseas flagship models shown in the image. The post also explicitly states that both the screenshot’s source and final scores require official confirmation; therefore, these numbers can only be considered as pre-release clues, not formal evaluations or conclusions about product capabilities. If the official version can replicate these results, the impact will primarily be on the competitive positioning of domestic models in the high-end market for Coding Agents and General Agents.

Sources:

The Value of Travel Agents Begins to Shift from “Planning for You” to “Executing Operations for You,” But Effectiveness Depends on Industry System Integration

Regarding the new “Feizhu Bangbang” feature in the Fliggy App, a blogger pointed out that its focus is not on generating a travel guide, but on connecting intent understanding to real business processes: users can make a one-sentence request for flight changes. The product also covers operations like check-in seat selection, hotel room upgrades, airport transfers, filling out immigration cards, and generating invoices for reimbursement, supporting parallel multi-tasking. This type of product addresses the issue where AI still leaves the execution cost to the user after organizing information. However, steps like flight changes, hotel upgrades, and invoicing depend respectively on GDS, hotel PMS, airline interfaces, and real order data. Whether they can be completed depends on whether the supply side’s interfaces and SOPs are prepared for the Agent. The visible evidence is product observation and analysis, not an indication that all scenarios are already stably available.

Sources:

Stats: Timeline Scans=600 Number of Bloggers Matched=53 Total Tweets Matched=298 Weighted Tweet Score=233.8 Original Tweets=121 RT Tweets=60 Crawl Attempts=4 Boundary Coverage Status=tail_confidently_crossed_target_boundary