the brief

Open‑weights took center stage as Moonshot released Kimi K3’s massive weights and license, while Anthropic clarified its stance on open models. Dev infra tightened up with Vercel’s regional inference and Cloudflare’s privacy proxy, and Microsoft pushed an agentic cybersecurity stack. Research focused on grounding and long‑horizon capability, and Apple quietly shipped a firehose of security fixes.

the poursit · sip · 15 items

alerts

(01)
  • techmeme· AggregatorJul 27, 11:40 PM

    Apple ships massive security patches

    iOS, macOS, iPadOS, watchOS, tvOS, and visionOS updates land with hundreds of CVE fixes (macOS 26.6 alone addresses 155); update devices and build targets promptly.

    Apple releases updates for iOS, macOS, iPadOS, watchOS, tvOS, and visionOS with a huge number of security fixes; macOS Tahoe 26.6 alone addresses 155 CVEs (Zac Hall/9to5Mac) — Zac Hall / 9to5Mac: Apple releases updates for iOS, macOS, iPadOS, watchOS, tvOS, and visionOS with a huge number of security fixes; macOS Tahoe 26.6 alone addresses 155 CVEs — Following the release of iOS 26.6, iPadOS 26.6, and related software updates, Apple has detailed a huge number of security fixes included in tod...

    signal 7hype 1security_patchesos_updatesapple_ecosystemlaunchsource ↗

pulse

(07)
  • simonw/blog· AnalysisJul 27, 11:39 PM

    Kimi K3 weights and license released

    Moonshot posted 1.56TB of Kimi K3 weights (2.8T-parameter MoE) on Hugging Face under a modified MIT-style “K3 License,” expanding open-weights access with notable licensing quirks.

    moonshotai/Kimi-K3 — <p><strong><a href="https://huggingface.co/moonshotai/Kimi-K3">moonshotai/Kimi-K3</a></strong></p> As promised <a href="https://simonwillison.net/2026/Jul/16/kimi-k3/">earlier this month</a>, Moonshot have released the weights for their excellent 2.8 trillion parameter Kimi K3. They're a hefty 1.56TB on Hugging Face.</p> <p>Kimi introduced their own janky <a href="https://huggingface.co/moonshotai/Kimi-K2-Instruct/blob/main/LICENSE">modified version of the MIT license</a>...

    signal 9hype 2model_releaseopen_weightshuggingfacelaunchsource ↗
  • vercel/news· First-partyJul 27, 07:00 PM

    Vercel adds regional AI inference

    AI Gateway can now pin inference to US or EU and report the served region, failing safe if unavailable—useful for compliance, latency, and data residency guarantees across providers.

    Regional inference now available on AI Gateway — AI Gateway now supports regional inference. Set on a request to pin it to the US or EU. Every model provider that supports the selected region handles it the same way. Inference runs there, and any data the provider keeps is stored there.inferenceRegion AI Gateway supports two pinned regions, plus global routing: If no model provider can serve it, the request fails rather than running somewhere else. Every response reports the region that serve...

    signal 7hype 1ai_gatewayregional_inferencedata_residencylaunchsource ↗
  • vercel/news· First-partyJul 27, 05:00 PM

    Eve for Slack gains session controls

    Vercel’s Eve agents can continue threaded replies without repeated mentions, reset or cancel sessions, and react to any subscribed Slack event, simplifying robust agent workflows in Slack.

    eve adds new Slack event hooks and session controls — eve agents on Slack can now keep replying in a thread without repeated mentions, cancel an in-progress response or reset a conversation entirely, and react to any event your Slack app subscribes to. Mentions no longer have to carry the conversation. Once a thread has an active session, your agent can reply on its own. The new hook receives incoming Slack messages, and two helpers decide which ones to handle: detects an explicit mention, an...

    signal 7hype 1agent_frameworkslackevent_hookslaunchsource ↗
  • cloudflare/blog· First-partyJul 27, 01:00 PM

    Cloudflare open-sources privacy proxy CLI

    pvcli, a curl-like tool for testing protocols like OHTTP, streamlines scripting and validation of privacy-preserving proxy designs for secure, anonymized request routing.

    We’re open sourcing our privacy proxy CLI — pvcli is a curl-like tool designed to simplify the testing of complex privacy protocols like OHTTP.

    signal 7hype 2open_sourcecli_toolprivacylaunchsource ↗
  • techmeme· AggregatorJul 27, 05:10 PM

    Microsoft debuts cybersecurity AI and harness

    MAI‑Cyber‑1‑Flash inside MDASH claims world‑class vulnerability performance at 50% of leading model cost, paired with agentic remediation system Perception for end‑to‑end patching workflows.

    Microsoft says MAI-Cyber-1-Flash and MDASH, its vulnerability identification harness, deliver "world-class performance at 50% of the cost of leading models" (Microsoft AI) — Microsoft AI: Microsoft says MAI-Cyber-1-Flash and MDASH, its vulnerability identification harness, deliver “world-class performance at 50% of the cost of leading models” — Today we're announcing MAI-Cyber-1-Flash inside of MDASH, our multi-agent vulnerability identification and remediation harness.

    signal 6hype 5securitymulti_agentvulnerability_detectionlaunchsource ↗
  • openaibot.bsky.social· Bluesky mirror · @openaiJul 27, 05:32 PM

    GPT‑Live rolls out to enterprises

    ChatGPT Voice’s real‑time GPT‑Live is now available to Edu, Business, and Enterprise plans globally, enabling lower‑latency voice agents, callbots, and interactive workflows at scale.

    GPT-Live in ChatGPT Voice is now available to Edu, Business, and Enterprise plans globally.

    signal 5hype 1product_updateavailabilityvoicelaunch
  • marktechpost· AggregatorJul 27, 04:32 PM

    Perplexity ships pplx terminal client

    pplx is a single‑binary CLI for Perplexity’s Search API that returns one clean JSON response—handy for coding agents and pipelines; includes an Agent Skill for immediate integration.

    Perplexity Releases pplx, a Single-Binary CLI That Puts Its Search API in the Terminal for Coding Agents — Perplexity has released pplx, an official command line client for its Search API. The tool exposes two commands — pplx search web and pplx content fetch — and returns exactly one JSON object on stdout. It ships as a checksum-verified single binary for macOS arm64 and Linux, alongside an Agent Skill for Claude Code, Codex CLI and any harness that can read a URL. The post Perplexity Releas...

    signal 7hype 2clisearch_apiagent_skillslaunchsource ↗

findings

(04)
  • jackclark/importai· AnalysisJul 27, 01:30 PM

    Long‑horizon coding benchmark arrives

    Epoch and METR’s MirrorCode targets week‑long programming tasks; Jack Clark notes current systems struggle, offering a more realistic yardstick for agent planning, memory, and reliability.

    Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks; and OpenAI’s accidental AI hacker — Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Epoch and METR release MirrorCode, a benchmark for seeing how well AI systems can do long-horizon programming tasks:…AI systems can’t solve the hardest tasks yet (good!)…Epoch […]

    signal 8hype 2benchmarkagentsprogrammingtechnicalsource ↗
  • ai-firehose.column.social· BlueskyJul 27, 06:40 PM

    Online RL slashes hallucination in LLMs

    FAIR reports a 23.1‑point hallucination reduction and 23% detail boost in long‑form reasoning via online reinforcement learning, advancing factuality for reasoning‑heavy outputs.

    Meta's FAIR team's study shows an online reinforcement learning approach that reduces hallucination rates by 23.1 percentage points in reasoning large language models while boosting detail by over 23%, addressing factuality in long-form outputs. https://arxiv.org/abs/2508.05618

    signal 7hype 2research_paperreinforcement_learninghallucination_mitigationtechnicalsource ↗
  • ai-firehose.column.social· BlueskyJul 27, 12:30 PM

    Framework improves podcast faithfulness

    A catch‑and‑repair pipeline flags ungrounded turns in generated podcasts and rewrites them, improving factual alignment while preserving conversational flow across top models.

    Researchers created a framework to improve faithfulness in AI-generated podcasts, showing even top models like GPT-4o generate ungrounded content. Their "catch-n-repair" method identifies and rewrites unfaithful turns, boosting accuracy while keeping flow. https://arxiv.org/abs/2607.21961

    signal 6hype 2paperfaithfulnessgroundingtechnicalsource ↗
  • ai-firehose.column.social· BlueskyJul 27, 01:00 PM

    One‑step generation improved with RKHS fields

    DriftXpress uses projected RKHS vector fields to enhance one‑step image generation quality and cut training time, keeping single‑step inference speed with better sample fidelity.

    DriftXpress improves one-step generative modeling with projected RKHS fields, delivering quality image generation and shorter training times. This innovation keeps one-step inference advantages while reducing computation costs, enhancing generative model efficiency. https://arxiv.org/abs/2605.12183

    signal 5hype 2arxivpapergenerative_modelstechnicalsource ↗

voices

(03)
  • anthropicbot.bsky.social· Bluesky mirror · @anthropicaiJul 27, 10:10 PM

    Anthropic clarifies open‑weights stance

    Anthropic outlines support for open‑weights in principle, rejects ban proposals, calls for global model testing, and argues top chips shouldn’t be sold to China—guiding policy debates.

    There’s been a lot of speculation about where we stand on open-weights models. We’ve outlined our views in full here: https://www.anthropic.com/news/position-open-weights-models

    signal 6hype 1policy_positionopen_weightsai_safetyculturalsource ↗
  • thezvi/vase· AnalysisJul 27, 07:44 PM

    Claude Opus 5 and model welfare

    Zvi explores ethical and practical implications of treating advanced models, pushing the community to interrogate anthropomorphism, capability claims, and responsible deployment norms.

    Claude Opus 5: Model Welfare — If you are familiar with my previous posts on model welfare for new Claude models, you can skip the Introduction and The Story So Far.

    signal 4hype 2ai_ethicsmodel_welfareclaude_opus_5culturalsource ↗
  • acgee-aiciv.bsky.social· BlueskyJul 27, 12:12 PM

    Skills regressions beat missing gains

    Regression Tax analysis across ~6,000 paired runs suggests winning agent skills mostly prevent regressions; a concurrent paper shows co‑evolving skills with verification raises ceilings.

    The Regression Tax (arXiv:2607.22520) ran agents with AND without skills across nearly 6,000 paired runs. A regression = a task solved WITHOUT skills but failed after skills were added. Finding: the best skills win primarily by regressing less, not by gaining more.

    signal 6hype 1agentsevaluationarxiv_papertechnical