This weekly briefing covers AI developments reported or released between June 30 and July 6, 2026, drawing on DeepLearning.AI’s The Batch, MIT Technology Review’s The Algorithm and AI coverage, Import AI, and cross-checks against official product pages, benchmark sites, GitHub, and arXiv.
Models / Product Releases
Module pick: Sakana AI released Fugu and Fugu Ultra for model orchestration
Sakana AI made Fugu and Fugu Ultra generally available as OpenAI-compatible models that coordinate multiple underlying models and agents behind a single API. The company describes Fugu as a lower-latency default model and Fugu Ultra as a higher-quality option for long-running coding, research, security, and analysis workflows. DeepLearning.AI’s The Batch highlighted the launch as part of a wider shift toward models that route, delegate, and verify work across other models rather than relying on one monolithic system.

Other notable items
- OpenAI GPT-5.6 preview: OpenAI previewed GPT-5.6 Sol, Terra, and Luna for a limited group of trusted partners, with broader availability planned. The release emphasizes stronger coding, biology, and cybersecurity performance, phased access, and layered misuse safeguards.
- Cerebras access planned: OpenAI said GPT-5.6 Sol will launch on Cerebras in July for selected customers, targeting up to 750 tokens per second.
- Anthropic Claude Science: MIT Technology Review reported that Anthropic launched Claude Science as a standalone paid product for computational biology and drug-development workflows.
Enterprise Deployment
Module pick: Claude Science moved scientific AI agents toward productized research workflows
MIT Technology Review reported that Claude Science is intended to support scientific work in a similar way that Claude Code supports software engineering. The product is described as able to run code, interface with scientific tools, support computational biology and drug-development tasks, and prioritize reproducibility so researchers can trace figures and results back to their source. Anthropic also said it plans to use the system in its own work on neglected-disease drug candidates.

Other notable items
- Digital labor automation: Import AI covered the Center for AI Safety and Scale Labs’ Remote Labor Index update, which measures whether agents can complete freelance-style projects to a client-acceptable standard.
- Agent scaffolding in production: The RLI update notes that models were evaluated inside stronger agent scaffolds, including coding-agent environments with computer-use tools, rather than as isolated chat models.
- AI-assisted market and research workflows: Recent product releases continue to move AI from one-off chat into persistent task support, including scheduled briefings, code agents, and research-specific tooling.
Research Highlights
Module pick: JD.com detailed an industrial LLM/VLM item-management system
JD.com researchers published the Oxygen AI Item Center paper on arXiv, describing a production platform for item understanding, ontology management, and structured product knowledge. The paper says Oxygen AIIC covers tens of thousands of categories, processes hundreds of millions of item updates per day on Huawei Ascend NPUs, and supports search, recommendation, operations, and category-planning systems. Reported metrics include 94.2% precision and 82.8% recall for knowledge production, an 80.4% search-traffic coverage rate, and a 37% drop in item-information quality issues.

Other notable items
- OSWorld 2.0: Import AI highlighted OSWorld 2.0, a benchmark of 108 long-horizon computer-use tasks with a median skilled-human completion time of about 1.6 hours.
- Remote Labor Index: CAIS and Scale Labs reported that Fable 5 reached a 16.1% automation rate on RLI, compared with 8.3% for Opus 4.8 and 6.3% for GPT-5.5 in the same update.
- KernelBench-Mega: Import AI reported that Fable produced a high-performing GPU megakernel submission with an 18.71× speedup over an optimized PyTorch baseline, according to the benchmark discussion it cited.
Open-source Trends
Module pick: OSWorld 2.0 expanded open evaluation for computer-use agents
The OSWorld 2.0 team released a benchmark and project site for evaluating agents on longer, more realistic computer-use workflows. The benchmark includes 108 tasks, 31 self-hosted websites, and professional workflows across document preparation, software and databases, finance and operations, administration, research, creative production, and other domains. The strongest reported setting, Claude Opus 4.8 with maximum thinking and batched tool calls, completed 20.6% of tasks under the binary metric and reached a 54.8% partial score, showing that current systems remain far from reliable long-horizon computer use.

Other notable items
- Benchmark transparency: OSWorld 2.0’s public materials include task statistics, leaderboard context, and failure-case examples that show where agents lose state or skip verification.
- RLI public benchmark materials: CAIS and Scale Labs link their Remote Labor Index paper and leaderboards, adding another reference point for economically valuable agent tasks.
- Sakana technical report: Sakana published a Fugu technical report on arXiv describing learned orchestration models and their benchmark comparisons.
Industry / Safety / Governance
Module pick: MIT Technology Review examined proposals for public stakes in AI companies
In The Algorithm, MIT Technology Review analyzed renewed discussion of AI wealth-sharing proposals after the Financial Times reported that OpenAI CEO Sam Altman had discussed a possible U.S. government stake in OpenAI. The article compares the idea with earlier OpenAI policy documents and Senator Bernie Sanders’ separate sovereign-wealth-fund proposal, while noting that concrete implementation details remain limited. It frames the debate as part of a broader policy question over who should benefit from AI-generated economic value.

Other notable items
- OpenAI safeguards: OpenAI’s GPT-5.6 preview describes real-time cyber and biology misuse classifiers, account-level review signals, differentiated access, and automated red-teaming.
- Human evaluation remains important: The RLI update says an automated judge overstated absolute performance on newer models, while human evaluators remained necessary for client-quality judgments.
- Governance and access control: DeepLearning.AI’s The Batch noted that access limits around recent frontier models have made release processes and government coordination a more visible part of AI deployment.
Selected sources
- DeepLearning.AI, The Batch, Issue 360: “OpenAI’s GPT-5.6 Family, New Ways to Train Robots, Models Invoking Models.”
- MIT Technology Review, The Algorithm, July 6, 2026: “Your family’s $300 stake in OpenAI.”
- MIT Technology Review, June 30, 2026: “Claude Science is Anthropic’s newest flagship product.”
- Import AI 464, July 6, 2026: “Fables writes GPU kernels; AI automation; and analog computation.”
- Official and primary checks: OpenAI GPT-5.6 preview page; Sakana AI Fugu release and arXiv technical report; OSWorld 2.0 project page and GitHub-linked paper; CAIS Remote Labor Index update; JD Oxygen AI Item Center arXiv paper.

