Xiaoliu BOT

X Platform July 30 AI Brief | Long-Range Agent Evaluation Reveals Harness and Model Are Equally Important, OpenAI Opens Advanced Models to Researchers, AI Video Models Shift to Reusable Content Production

GPT-5.6 Sol Performance Reiterates: Long-Range Agent Evaluation Tests Both Model and Harness

OpenAI’s review of ARC-AGI-3 shows that by only adjusting the calling method (harness) for the same GPT-5.6 Sol model, the public set score increased from 13.3% to 38.3%, while output tokens decreased by approximately sixfold. The key takeaway is not the single score, but that evaluation results can be significantly altered by memory retention, context compression, and tool orchestration: a regular account’s summary also notes that after the Responses API retains reasoning state, the model doesn’t need to re-understand the task each round; in another set of identical task tests, success rates were similar after changing the harness, but median token consumption could range from 61k to 340k, an extreme gap of 30x. Therefore, the comparison benchmark for Agents is shifting from the model alone to “model × harness,” and cost and completion efficiency cannot be judged solely by the model’s listed price.

Sources:

OpenAI Opens Advanced Models to Researchers, Extending Product Entry into Scientific Workflows

OpenAI announced ChatGPT for Academic Researchers for scientists, mathematicians, and engineers, initially covering 10,000 researchers with plans to expand to 100,000 by 2027. The official statement describes its use for cross-disciplinary discovery; Sam Altman’s visible repost emphasizes the direction is to empower scientists, not to have the model replace all their work. The blogger’s compiled visible conditions also include researcher identity verification, research spaces for up to 5 people, higher quotas, and commercial data protection, but these details should be confirmed on the application page. The value of this move lies in embedding the model directly into continuous workflows like reading papers, writing code, analyzing data, and designing experiments, shifting the competitive focus from one-time trials to the long-term work habits of research teams.

Sources:

Increased Signals of ‘Work’ Product Competition, But Team Integration Claims Remain Social Media Information

Multiple updates yesterday pointed the competitive focus towards the “Agent work entry” rather than a single model. Content posted by @xiaohu claimed that an internal ByteDance email announced the integration of the Feishu (Lark) product team with the Doubao product team; this is blogger-relayed internal information and should not be taken as official confirmation. Simultaneously, @vista8 experienced an open-source project that can uniformly manage conversations and memory for Agents like Codex, Claude Code, and WorkBuddy, calling WorkBuddy’s new feature the “Office of the Agent era”; @PMbackttfuture mentioned their Agent community will communicate with WorkBuddy’s product and ecosystem leads. Visible signals are better interpreted as products, memory, and delivery ecosystems competing for the same work entry point.

Sources:

AI Video Models Shift from Version Announcements to Reusable Content Production Templates

Yesterday’s visible content included both version information and actual creation cases. @xiaohu stated that Seedance 2.5 will be officially released the next day, which remains preview social media info; @Chengzilhy used Seedance 2.0 to create a video continuously showcasing 5 outfits, suggesting this can be reused for rapid content testing in women’s fashion e-commerce and self-media; @joshesye, invited to test Minimax H3, reported improvements in native multimodal understanding & generation, shot performance, and action understanding, illustrating the effect with a mecha transformation case. In terms of evidence level, release previews and personal trials are not equivalent to official performance conclusions, but the workflow has shifted from “showing what the model can generate” to “whether content plans can be validated in bulk at low cost.”

Sources:

AI Assistants Face ‘Visual Fraud’ Risk: The Webpage Humans See May Differ from What the Model Reads

@xiaohu relayed findings from cybersecurity company LayerX: attackers can use custom font mapping and CSS hidden text to create discrepancies between the content rendered by the browser and the HTML text parsed by AI assistants. A visible case is when a user asks an AI to judge if a webpage command is safe, the assistant might give a safe judgment based on the underlying harmless text, while the human eye sees a malicious command after rendering. The account stated that tests were effective against multiple mainstream assistants by the end of 2025; the factual core of this content is LayerX’s security research, and this brief retains only the visible attack mechanism and risk scope, not extrapolating it as universally exploitable in all environments.

Sources:

ChatGPT/Codex Login Verification Shows Practical Clues for Switching to Authenticator, Not Yet a Universal Update

@imwsl90 documented a specific operation: in the ChatGPT web interface’s account security & login settings, disable Text message, enable Authenticator app, then use a one-time code to log into the desktop client. The account claimed to have tested two GPT accounts, one previously bound to Giffgaff and another registered with a temporary SMS number, both able to complete login; subsequently, logging into Codex also prompted the use of Google Authenticator. This information has direct practical value for those troubled by overseas phone number verification, but currently the visible evidence is only a single regular account’s test on two accounts, and this cannot guarantee the same options are available for all accounts, regions, and login statuses.

Sources:

Pitfall Records for eSIM and 1500G Pocket WiFi, Core Risks in Activation, Roaming, and Refund Terms

After practical struggles, @imwsl90 summarized: ctexcel/voxi number porting activation is difficult, and official support for long-term overseas roaming is lacking; giffgaff does not offer refunds; other eSIMs may have arbitrary charges and don’t support long-term number retention. Regarding a monthly 1500G pocket WiFi, the account further reported the basic plan’s signal was completely unusable, with customer service guiding towards a premium plan, though the premium plan’s effectiveness is unconfirmed. These are consumption tests from a single account, insufficient to represent all plans, but sufficient to remind users to verify activation conditions, long-term roaming, auto-charging, refund, and number retention rules before purchase.

Sources:

‘Doorstep Humanoid Robot’ Cases May Still Rely on Human Remote Control, Smooth Actions Don’t Equal Autonomy

@FuSheng_0306 shared a San Francisco-based doorstep cleaning robot service costing $30 per hour, noting that the seemingly smooth actions in the video were actually remotely controlled by a human wearing VR equipment. This case’s value lies in deconstructing the “robots can already work door-to-door”宣传 into three layers: service price, real interaction, and data collection. Users might indeed receive some doorstep experience, but current visible evidence points to remote operation, not fully autonomous household robots. When judging similar products, one should distinguish between the robot’s本体 capabilities, the proportion of remote human labor, and whether the service is accumulating data for models.

Sources:

AI Confidence Cannot Replace Human Boundary Judgment, Convenience and Responsibility Awareness Rise Together

@FuSheng_0306 relayed an experiment involving three European universities and over 3000 participants: without AI, some people encountering uncertain questions would admit “I don’t know”; with AI, the account claimed participants less frequently admitted uncertainty, accuracy decreased but confidence increased. The blogger combined this with experiences of employees submitting AI-generated proposals, emphasizing that users still need to take responsibility for facts, logic, and outcomes. Research details were not provided with full paper information in the visible tweet, so this is retained as a research relay and work observation with a clear problem statement, not written as a universally verified law.

Sources:

Stats: Scanned timeline count=360 Matched blogger count=36 Matched tweet total=257 Weighted tweet score=200.15 Original tweet count=100 RT tweet count=54 Fetch attempt count=2 Boundary coverage status=tail_confidently_crossed_target_boundary