Disclosure: TechSifted has no affiliate relationship with OpenAI, Google, or Sierra. This is editorial news coverage.

Three things happened in AI this week that are worth your attention. None of them are hype cycles. One is a genuine surprise, one is an infrastructure move that could matter a lot, and one is a useful tool that comes with a catch you should know about.

OpenAI Published 719 Math Manuscripts on GitHub

On October 6, OpenAI dropped a public GitHub repository containing 722 mathematical manuscripts produced by an unreleased internal model (Source). The papers are organized into 372 result families, each classified by mathematical discipline.

The project came out of model evaluation work. According to OpenAI, performance on its existing math evaluations had saturated, so it posed the model roughly 4,000 problems and kept the results that met a significance bar. On average, each result used about three hours of ChatGPT Pro thinking compute. The repository includes PDFs, source files, and Lean formal proofs for approximately 42% of the top-line results. The license is Apache-2.0.

The claimed results are striking. The collection includes work on a zero-free region for the Riemann zeta function and a proof of the Hodge Conjecture for CM abelian varieties. Those are serious problems. OpenAI lists both as exceptions to its standard procedure, and says the Riemann zeta writeup was human edited for readability. OpenAI is careful to note that not every manuscript has a Lean formalization and that some unformalized results “could have issues.”

There’s a real tension in how this was released. MIT mathematician Andrew Sutherland told Scientific American that any claims about one-shotting problems with a single agent should be treated as unverified until OpenAI releases the model and people can replicate the results (Source). The release also leaves out the prompts that an advisory group at the Institute for Advanced Study recommended disclosing. The point isn’t that the results are wrong. It’s that announcing “we proved this” before the field has reviewed it is a strange way to behave if you want to be taken seriously as a research contributor. Putting the full repository in public, where anyone can inspect the proofs, does count for something.

As someone who used to write code that had to be verified by someone else before it shipped, the analogy holds: the Lean formalizations are the machine-checked tests. The 58% of top-line results without formalizations are the modules with manual review pending. Worth watching, not celebrating yet.

Find out more: github.com/openai/math

Sierra and Meta Want a Standard for How Your AI Agent Shops

On the same day, Sierra and Meta announced the Personal Agent Protocol, an open standard that tries to define how a personal AI agent authenticates with a business and takes actions on your behalf.

The founding partners read like a retail and payments roll call: Walmart, Shopify, Stripe, Rocket Companies, Genesys, Instinct. The idea is that right now, if you ask your AI agent to check your order status on a retailer’s site, it walks through a website like a person. Clicks, forms, session state. That’s slow, fragile, and gives the business no visibility into what’s happening. The protocol proposes a structured alternative.

The mechanics are OAuth-based. An agent connects through a company’s website, its APIs (MCP and OpenAPI are explicitly mentioned), or directly with a business’s own agent. A session can start as a guest. The agent can check stock or look up return policies without logging you in. When a task needs your account, you sign in, and you decide whether the agent gets read-only or write access.

A v0.1 spec is expected later this month, followed by design workshops and a reference implementation. This is still early. The spec isn’t out yet and “founding partners” on a press release is not the same as shipping integrations.

The protocol that wins here will look like OAuth did for web apps: invisible when it works, painful when it doesn’t. The interesting question is whether this one gains enough adoption to become the standard or whether it becomes one of five competing standards. The MCP alignment is a good sign.

Find out more: sierra.ai/blog/introducing-personal-agent-protocol

Google’s SynthID Detector Is Now for Everyone

On October 7, Google opened its SynthID Detector to the public at synthid.com. The tool checks images, videos, and audio files for AI watermarks. It works across content from Google, OpenAI, NVIDIA, and Kakao. Apple is expected to be added soon.

The scale of SynthID’s deployment is worth noting. Google says it has watermarked more than 180 billion images and videos and the equivalent of 240,000 years of audio since the technology launched in 2023. The Detector now gives anyone a direct way to check whether a specific file carries one of those watermarks.

Here’s the part the headline usually buries: this is explicitly not a general AI detector. The tool’s own page says so directly: “This is not a general AI detector.” If a file was made with a system that hasn’t adopted SynthID, the Detector won’t flag it. That covers a lot of content. Claude-generated images, local open-source image models, tools outside the current partnership group. None of those would trigger a positive result. Absence of a watermark is not evidence of authenticity.

The practical value here is narrow. Did this image come from a Google, OpenAI, NVIDIA or Kakao tool? That’s the question it answers. Journalists verifying viral images, brands checking for unauthorized AI-generated content with their likenesses. As a broad AI content detection tool, it isn’t one.

Find out more: Google SynthID Detector launch coverage