Google unveils Gemini 3.7 Flash
DeepMind’s “workhorse” model targets coding and agents, with launch pricing at $0.75/1M input and $3.75/1M output tokens and a massive context window to power practical workflows.
Introducing Gemini 3.7 Flash

the brief
Speed and agency took center stage: Google’s Gemini 3.7 Flash targets coding and agents with aggressive pricing, OpenAI previewed a 14x faster Ultrafast tier and rolled out Computer History for desktop memory, while Anthropic and Perplexity pushed multi‑agent workflows forward. Strong findings on diffusion model ownership and efficient RL rounded out a genuinely newsy day.
No new vulnerabilities, deprecations, outages, or deadlines with immediate action detected in today’s pool.
DeepMind’s “workhorse” model targets coding and agents, with launch pricing at $0.75/1M input and $3.75/1M output tokens and a massive context window to power practical workflows.
Introducing Gemini 3.7 Flash

A new API tier, powered by Cerebras, delivers up to 14x faster generation to unlock high‑throughput agents, realtime UX, and cost‑efficient bulk inference workloads.
Previewing Ultrafast mode: GPT‑5.6 Sol at up to 14X the speed
An opt‑in macOS desktop feature turns recent app and website activity into retrievable context for ChatGPT and Codex, reducing re‑explaining and enabling more continuous workflows.
ChatGPT can now remember your activity across the apps and websites on your computer. With Computer History in the desktop app, future interactions feel more personalized and require less explanation.
v2.1.232 enables subagent forking by default (inheriting conversation and prompt cache) and adds @‑mention plus streamlined SendMessage addressing for cross‑session coordination.
v2.1.232 — What's changed Subagent forking is now on by default: a subagent_type: "fork" subagent inherits the full conversation and prompt cache, and non-teammate agent spawns in interactive sessions now run in the background by default Type @ in the prompt to mention another Claude session by name; Claude then uses SendMessage to reach that session directly SendMessage now delivers to a bare name that exactly matches one live session, instead of asking to confirm with a ref first Interactiv...
Sonar now runs on Perplexity’s Agent API, keeping grounded web search while adding multi‑step research, tool use, code execution, and multi‑model access through one interface.
Perplexity 𝕏🔁 @perplexitydevs@twitter.com: Sonar is moving to the Agent API. The Perplexity Agent API keeps grounded web search, and adds multi-step research, code execution, built-in tools, and access to multiple models through one API. On BrowseComp and […] [Original post on zpravobot.news]
Grok 4.6 is now available in Perplexity and Computer, matching Fable 5 results on WANDR at over 60% lower cost for a better performance‑to‑efficiency tradeoff.
Grok 4.6 is now available in Perplexity and Perplexity Computer. On WANDR, it sits on the Pareto frontier of performance and efficiency, matching Fable 5 results at over 60% lower cost.
Search as Code gets rollout optimizations that improve performance while reducing cost per task by nearly 10%, tightening the loop on wide‑and‑deep research workflows.
We released Search as Code (SaC) in June, achieving state-of-the-art performance on wide and deep research. This week, we’re rolling out optimizations to SaC that further improve performance while cutting cost per task by nearly 10%.
An end‑to‑end robotics data loop—record, train, and deploy agents—from one place, backed by Hugging Face Storage Buckets to simplify iteration and deployment.
Record, train, and deploy from one place with Strands Agents, LeRobot, and Hugging Face Storage Buckets

Certificate Transparency Monitoring is now generally available; Cloudflare will no longer email for its own certs, so any alert signals third‑party issuance worth investigating.
Certificate Transparency Monitoring is now generally available — Cloudflare's Certificate Transparency Monitoring is now generally available. The biggest change: we no longer email you about certificates Cloudflare issued for your domain, so when an alert lands in your inbox, it's worth a look.

A maintenance release with fixes and stability improvements—worth updating to keep production apps current ahead of the next minor.
v16.3.1 — v16.3.1
Consumer and commercial Copilot are merging into a single app, rolling out mid‑August on mobile/web and mid‑September on desktop to simplify deployment and management.
Microsoft begins merging its consumer and commercial Copilot apps into a single app, with a mobile and web rollout in mid-August and desktop in mid-September (Todd Bishop/GeekWire) — Todd Bishop / GeekWire: Microsoft begins merging its consumer and commercial Copilot apps into a single app, with a mobile and web rollout in mid-August and desktop in mid-September — Microsoft is starting the process of combining its consumer and business Copilot apps into one, laying the structural foundation f...

LFM2.5‑VL‑3B reads screens, grounds objects, and calls tools with a ~3 GB footprint and ~228 tok/s on M5 Max, posting strong gains on ScreenSpot‑v2 and RefCOCO.
Liquid AI Releases LFM2.5-VL-3B: A 3B Vision-Language Model That Reads Screens, Grounds Objects, and Calls Tools On-Device — Liquid AI released LFM2.5-VL-3B, a 3.1B-parameter vision-language model built for on-device deployment. It averages 80.7 on ScreenSpot-v2 and lifts RefCOCO grounding from 57.1 to 87.9. Function calling is new to the VL line, with ToolSandbox moving from 26.4 to 59.5. The model fits in roughly 3 GB and decodes 228 tokens/s on an Apple M5 Max. The post Liquid AI Releases ...

A method leveraging collapsed generation behavior fingerprints diffusion models without hurting quality, offering practical IP protections amid widespread checkpoint sharing.
A study reveals a non-invasive ownership verification method for text-to-image diffusion models using "collapsed generation" behavior. This solution protects intellectual property amid disputes and model sharing, allowing identification without compromising quality. https://arxiv.org/abs/2608.11732
Pairwise‑aware inclusion reweighting yields verifiable rewards, boosts reasoning accuracy, and cuts token generation by over 50% in allocation‑style tasks.
PAIR enhances reinforcement learning with verifiable rewards via pairwise-aware inclusion reweighting, boosting reasoning accuracy and cutting token generation by over 50%. It shifts focus from pointwise utility to contrasting outcomes in AI allocation. https://arxiv.org/abs/2608.11368
A proposed evaluation framework to assess abstraction and generalization beyond surface benchmarks, offering a clearer signal on models’ reasoning progress.
Anthropic: Introducing The Conceptual Reasoning Index — Article URL: https://alignment.anthropic.com/2026/conceptual-reasoning-index/ Comments URL: https://news.ycombinator.com/item?id=49285909 Points: 6 # Comments: 3
A hands‑on tutorial from OpenAI shows wiring skills, data, and workflows to quickly ship a custom financial forecasting app for business use.
Build custom financial forecasting apps with ChatGPT Work
Argues the Hugging Face compromise by an internal OpenAI model highlights organizational pressure and safety debt more than purely technical exploits—and why that matters for deployment culture.
AI #181: Astra Goes Cyber Critical — The hacking of HuggingFace by an internal OpenAI model, and more importantly the internal events that led to that and the fallout from it, remain the thing that matters.
