Category: Data governance

  • When the tool you trust becomes the competitor

    When the tool you trust becomes the competitor

    A proof made the headlines. The more durable story is that two sophisticated users spent a year putting unpublished work into a vendor’s tool without a verified answer to whether the vendor could read it, and almost every organisation is running the same unexamined arrangement.

    Key points

    • Two research mathematicians put a year of unpublished work through a vendor’s coding assistant without establishing whether the vendor could access it or train on it.
    • OpenAI denies that its researchers or agents saw the work, while stating in writing that it cannot rule out that de-identified data from the pair’s own product usage improved its models.
    • The governance failure is not the dispute. It is that the boundary went unverified for a year, by users with every reason and every capability to check it.
    • The same unverified boundary sits under most organisations’ use of AI coding and drafting tools, because governance reviews rarely treat developer tools as a data channel at all.

    On 8 September 2026, OpenAI announced that an unreleased internal model had produced a finite-time blowup proof for the three-dimensional Navier–Stokes equations, generated by roughly 10,000 concurrent agents over 88 hours and formalised in Lean.1

    1Fortune, 8 September 2026

    Two qualifications matter. The proof relies on a smooth external forcing term, a route the written Millennium Prize question permits but which most working mathematicians exclude from the problem they care about. And the Clay Mathematics Institute has not accepted the result: it still lists Navier–Stokes among its unsolved problems, its rules require publication and roughly two years of acceptance by the field before evaluation begins, and its president has said the assessment will be deliberately unhurried.2

    2Implicator.ai, 9 September 2026

    So the mathematics is unresolved and will stay that way for some time. That is not the interesting part for a governance audience.

    The dispute, briefly

    Tristan Buckmaster, a mathematics professor at NYU’s Courant Institute, and Levent Alpöge, a mathematician employed at Anthropic, had worked for roughly a year on related blowup results, as a personal collaboration with no institutional involvement from either employer. They used several assistants throughout, including Anthropic’s Claude and OpenAI’s Codex, and Buckmaster has stated that their Codex sessions held all of their drafts for the entire project.3

    3Tom’s Hardware, 9 September 2026

    On 3 September, Alpöge passed on rumours that their progress had reached OpenAI. Buckmaster went public on the night of 7 September, hours before OpenAI’s announcement.

    He says he asked directly whether the model had been given access to, or trained on, those Codex sessions. He was told the model did not look up user data, and says he received no answer to the training question.4

    4The Next Web

    OpenAI subsequently denied that its researchers or agents saw the pair’s work before publication, and Sébastien Bubeck has said the proofs differ substantially. In the same published statement, the company wrote that while unlikely, it could not rule out that “de-identified data derived from their usage of our products helped improve our models”.5

    5OpenAI statement, reported by Futurism

    That sentence is the one worth sitting with. It is not an allegation, and it is not about customers in general. It is a qualified answer about these two users’ own material: a boundary they assumed was clean turns out to depend on contract terms, product tier, and configuration choices that almost nobody audits.

    Why this is a governance story, not a maths story

    Two working mathematicians, sophisticated users by any definition, ran a year of unpublished, commercially and academically valuable intellectual property through a vendor’s coding tool. They did not have a clear, verified answer to a basic governance question: can the vendor’s own models see this, and can it be used to improve them? That question sat unresolved for a year, under a live research programme, until a dispute forced it into the open.

    Swap “unpublished proof” for “unreleased product roadmap”, “draft client contract” or “unfiled patent claim”, and it stops being a story about mathematicians and becomes a story about any organisation letting staff use a vendor’s AI coding or drafting tool on sensitive work. The question is identical: what do the vendor’s terms actually permit, and does anyone check, or does everyone assume the boundary holds because it would be inconvenient if it didn’t?

    The relevant test for a Swiss deployer

    Regulation (EU) 2024/1689, Art. 50, in force since 2 August 2026, requires transparency about AI-generated content and points more broadly toward a standard in which AI systems’ data handling is legible to the people using them, not merely defensible after the fact.

    For a Swiss organisation the primary instrument remains the Federal Act on Data Protection, with the AI Act reaching the organisation only through a genuine EU nexus. Under either, the practical question is not whether the vendor’s contract technically permits this, since it almost certainly does in some qualified form, but whether anyone has read the clause and whether staff behaviour matches what it says.

    Most organisations have not done this exercise for their coding assistants specifically. Data governance reviews focus on the obvious surfaces, email, file storage, CRM, and treat developer tools as infrastructure rather than as a channel through which proprietary and unpublished material routinely gets typed into a third-party system.

    What to actually check

    Three things, none of which require legal counsel to start:

    • What do the vendor’s terms say about training-data use for your specific product tier, and does your subscription match the tier the favourable language applies to?
    • Is there a technical or contractual distinction between “the vendor can access this” and “the vendor will train on this”, and which one does your configuration guarantee?
    • Would you know if the answer changed? Vendor terms shift, and the review that established your comfort level a year ago may not describe today’s product.

    A story still in motion

    The mathematical status of the claim, and the factual dispute between the company and the researchers, are both unresolved. Nothing above depends on how either turns out. The governance gap, a working boundary between using a tool and training on your unpublished work that most users have never verified, was exposed regardless.