Terence Tao — Fields Medal winner, UCLA professor, and by most accounts the greatest living mathematician — publicly shared a ChatGPT conversation in which he works through aspects of a potential counterexample to the Jacobian Conjecture, an unsolved problem that has resisted proof or disproof since 1939. The shared link hit the Hacker News front page with 544 upvotes and 339 comments, and the discussion reveals something genuinely useful for anyone doing complex analytical work with AI tools. But the biggest pitfall waiting here is the obvious misread: that this is a story about AI being smart enough to tackle hard mathematics. That's not what happened. The story is about something more transferable and more immediately useful — how an expert structures a session to make AI productive for problems where the model cannot be trusted as an authority.
Missing that distinction is what separates people who get real work done with AI from people who produce expensive-looking nonsense.
What the Jacobian Conjecture Is — and Why a Counterexample Would Be Significant
To make sense of what Tao was doing, you need enough context to understand why this particular problem matters. The Jacobian Conjecture doesn't require advanced mathematics to follow at a conceptual level.
Ott-Heinrich Keller posed the conjecture in 1939. The question is simple to state: given a polynomial map — a function that takes coordinates and outputs new coordinates using polynomials — from n-dimensional complex space to itself, if the Jacobian determinant (a specific matrix determinant formed from the partial derivatives of the map) is a nonzero constant everywhere, must the map be invertible with another polynomial map?
In two dimensions, this asks something that sounds intuitive: if you smoothly deform the plane in a polynomial way and the deformation has constant "volume-stretching" everywhere, can you always undo that deformation with another polynomial? Sounds like the answer should be yes. Mathematicians have been unable to prove or disprove it for 87 years.
The conjecture appears on Smale's canonical list of important open 21st-century problems. It has an ugly history of swallowing careers — at least five high-profile claimed proofs or counterexamples have circulated over the decades, each eventually found to be flawed. The mathematical community treats it as genuinely treacherous territory. Attempted proofs in dimensions above two have repeatedly failed at structural points that seemed minor until scrutinized.
A counterexample would be a specific polynomial map where the Jacobian determinant is constant but the map provably has no polynomial inverse. That would settle the conjecture by falsifying it. Finding such an object, if one exists, isn't an algebraic trick — it requires deep structural insight about how polynomial maps work globally.
What the shared conversation appears to show is Tao working through the properties of a specific candidate map that might satisfy the conditions for a counterexample. He is not asking ChatGPT whether the conjecture is true. The model would be unable to answer that question reliably and shouldn't be trusted to try. What he is doing is using the model as a structured surface for reasoning — articulating premises, checking logical consistency step by step, probing whether algebraic manipulations hold, identifying what a given line of argument actually requires to go through.
The role partition is everything. Tao supplies the mathematical judgment, the candidate object, and the critical verification of the model's output. ChatGPT supplies structured response — something to think against, something that catches when stated premises aren't being honored, something that offers alternative framings of the same logical situation. It is less like a research assistant and more like a very patient whiteboard that can write back.
That dynamic — expert judgment driving the session, AI acting as a disciplined sounding board — scales. It scales to consultants, product managers, developers, and agency strategists working through hard problems in their own domains. Understanding how the dynamic works at Tao's level is the entire point of analyzing this story.
Why This Matters Right Now
Twelve months ago, the mainstream narrative about AI and genuinely difficult intellectual work was roughly stable: useful for summaries, drafts, and code completion; unreliable for complex domain reasoning. That frame has been eroding throughout the past year as reasoning-focused model variants have become meaningfully better at structured multi-step inference.
But the more significant shift isn't capability — it's credibility signal. Tao is not the first serious researcher to use AI in exploratory mathematics. Papers citing AI-assisted proof exploration have appeared. Individual researchers have been quietly using models as thinking aids for some time. What is different about Tao explicitly publishing this conversation is the visibility and apparent endorsement. He used OpenAI's link-sharing feature, which means he made a deliberate choice to surface the conversation as a document — not buried in a footnote, not mentioned in passing. The act of sharing is itself a methodological statement.
That matters for organizations well beyond mathematics. One of the persistent barriers to AI adoption among domain experts isn't model capability — it's professional legitimacy. Lawyers, financial analysts, consultants, engineers: people in these roles often carry an implicit risk calculus around AI use in serious work. The Tao example doesn't fully neutralize that concern, but it shifts the default. It makes "I used AI to stress-test my reasoning" feel more defensible than it did.
There is also a genuine technical shift underneath the credibility story. Current frontier models — particularly the o-series reasoning models from OpenAI and comparable reasoning-focused variants from Anthropic — handle long-context analytical dialogues substantially better than their predecessors from 18 months ago. Maintaining logical consistency across a 30-message conversation, honoring premises set early in the session, tracking multiple interrelated constraints: these are things that broke down much more frequently in 2024-era models. The Tao conversation, from what's visible, runs long and stays coherent. That kind of session would have been notably less reliable in earlier model generations.
The institutional context has also matured. Major professional services firms and research organizations have spent the past year developing frameworks for responsible AI use in expert work. Those frameworks are now stable enough that individual practitioners can deploy AI tools within clearer guardrails. The infrastructure for serious AI use has caught up with the raw capabilities. What that means practically: teams that were waiting for institutional permission to use AI for complex analytical work now have more cover to move.
Practical Implications for Small Teams
The reason this story lands specifically for independent practitioners and small teams is that they're exactly the constituency who can benefit most from this workflow — and also the one most likely to misread what "benefit" means.
Scenario 1: The consultant stress-testing analysis before a client presentation.
A strategy consultant working solo has developed a market entry framework. She's been in the weeds for two weeks. The risk at this stage isn't obvious error — it's the subtle logical inconsistencies that accumulate when you're too close to the work to see them. The Tao workflow applied here: articulate the core argument explicitly at the start of an AI session ("my conclusion is X because Y because Z"), then ask the model to find the assumption doing the most weight-bearing that you haven't verified. Ask it to steelman the opposing position as strongly as possible. Ask it where your argument would collapse if a specific market condition turns out to be false.
The model won't necessarily know things the consultant doesn't. But it will expose the places where the argument relies on unstated premises — which is exactly the vulnerability that client questions will hit. Catching that before the presentation is worth the session time.
Scenario 2: A developer working through an architectural decision.
A solo developer needs to choose between two data architecture approaches for a SaaS product. One is simpler now but creates potential scaling problems; the other is more complex upfront but more flexible. Confirmation bias is dangerous here — you tend to build the case for whatever approach you're already leaning toward.
Using AI as a structured adversary: write out your currently preferred approach explicitly, then ask the model for the five most damaging objections to it. Then ask for the five most damaging objections to the alternative. Then — and this is the step most people skip — ask which set of objections seems harder to address and why. The model's assessment may not be correct, but the process forces you to articulate the decision landscape more honestly than internal deliberation alone tends to produce.
Scenario 3: An agency pitching in unfamiliar territory.
An agency team is preparing a pitch for a client in a sector they haven't worked in before. The strategy makes internal sense but carries obvious blind spots about market-specific dynamics. The AI session here functions as fast-turnaround domain-expansion — specifically, to identify which assumptions in the strategy are doing the most structural load-bearing. "Our proposed approach depends on the assumption that customer acquisition in this sector works via X. What would need to be true for that assumption to hold? What observable evidence might suggest it doesn't?" This generates the specific questions that need independent research — not the answers, but the questions. That preparation changes the pitch meeting, and often surfaces a conversation starter the client finds unexpectedly sharp.
Scenario 4: A founder working through a high-stakes pivot decision.
A solo founder has data pointing in two directions and strong intuitions that may or may not be reliable. The challenge isn't finding more data — it's being honest about which data is genuine evidence and which is confirmation of what they already want to do. Loading the decision into an AI session and forcing explicit articulation of both cases — FOR staying the course and FOR pivoting — then having the model probe the reasoning structure, often surfaces the one thing being avoided. What is the single key uncertainty that, if resolved against you, would change the decision? Getting clear on that question before a major pivot can save months.
What these scenarios share is the core dynamic: human judgment in charge, AI as a disciplined interlocutor. The model challenges, the human verifies and decides.
How to Actually Run This Workflow
The mechanics matter more than most people think. Most AI conversations look like enhanced search queries — a question, an answer, done. The Tao-style reasoning session is structurally different: an iterative dialogue where each exchange builds on an established logical scaffold, and where the human is actively verifying output rather than accepting it.
Start every session with explicit premise-setting. Don't open with a question. Open by stating your current state of knowledge, your working assumption, and the specific thing you're trying to resolve. The model needs the constraint space established before it can be useful as an interlocutor. Something like: "I'm working through a decision in [domain]. Here's the context. Here's my current hypothesis. Here's what I'm uncertain about." That framing shapes everything that follows.
Ask for interrogation, not confirmation. "What's wrong with this reasoning?" produces more value than "Is this reasoning correct?" The former invites adversarial probing; the latter invites a confident affirmation that may not be earned. This is one of those things that seems obvious but almost nobody does naturally — the instinct is to present your best case and ask if it's good. The productive move is to present your best case and ask how it fails.
Maintain the logical thread deliberately. In sessions that run long, periodically ask the model to restate the framework as it currently understands it. This catches drift — places where the conversation has implicitly shifted a premise without explicitly noting it. In Tao's mathematical work, this would be critical for catching algebraic errors of omission. In business analysis, it catches the moment when "assumption X holds approximately" has quietly become "assumption X is treated as established."
Treat the model's output as candidate material, not conclusions. When it says "this seems to imply Y," that's a prompt to check whether Y actually follows — not a verified result. The expert judgment comes in at verification, not at generation. That division is what makes the workflow function. What tripped up early practitioners of this approach (and this is a pattern we've noticed across how people describe their AI workflows) is treating the model's fluent output as trustworthy because it sounds right. Fluency and correctness are orthogonal in LLMs, and the gap between them is exactly where errors hide.
For tool selection: ChatGPT's o-series reasoning models are what Tao used, and they're specifically designed for extended logical reasoning chains. Claude Sonnet and Opus offer a different reasoning texture that many practitioners find better for very long sessions where subtle contextual consistency matters. A multi-tool approach — reasoning structure with a frontier LLM, computational verification with Wolfram Alpha, factual grounding with Perplexity — is more reliable than any single model.
Comparison: AI Tools for Complex Reasoning Work
| Tool | Best for | Free plan | Starting price | Key differentiator |
|---|---|---|---|---|
| ChatGPT (o3/o4 reasoning) | Long reasoning chains, math, formal logic | Yes | ~$20/mo (Plus) | o-series models built for extended inference; what Tao used |
| Claude (Sonnet/Opus) | Long-context analysis, nuanced consistency | Yes | ~$20/mo (Pro) | Better at holding subtle context across very long sessions |
| Gemini Advanced | Deep integration with Google Workspace | Yes | ~$20/mo (Advanced) | Best option if your work lives in Docs, Sheets, Drive |
| Perplexity Pro | Factual research with source grounding | Yes | ~$20/mo (Pro) | Web-grounded responses significantly reduce hallucination on verifiable claims |
| Wolfram Alpha Pro | Mathematical and computational verification | Yes | ~$7/mo (Pro) | Doesn't reason, but calculates reliably — trustworthy for checking specific outputs |
| NotebookLM | Document-grounded analytical dialogue | Yes | Free / ~$20/mo (Plus) | Anchors responses to your uploaded documents; minimal hallucination within that scope |
Our take: for reasoning-partner work, the meaningful choice is between ChatGPT's o-series (stronger on formal logic and mathematical structure) and Claude (stronger on maintaining contextual consistency across long conversations). They're different enough that for high-stakes analytical work, having access to both is worth the combined subscription cost. The other tools in the table fill specific verification roles rather than serving as primary reasoning partners.
What the HN Community Is Saying
The 339-comment thread follows the standard HN distribution: a confident skeptical contingent, a smaller optimistic contingent, and a larger group of practitioners quietly describing what they actually do. The signal-to-noise ratio is better than average.
The skeptics advanced two main arguments. The first is that ChatGPT is unreliable for serious mathematics — it produces fluent-looking algebraic manipulations that contain subtle errors, and a less-expert practitioner could be led badly astray by output that sounds authoritative. This is a legitimate and important concern. The counter, articulated by several commenters, is that Tao clearly isn't using the model to produce authoritative mathematical output. He's using it as a structured sounding board and verifying everything against his own reasoning. The model's outputs are prompts for verification, not results.
The second skeptical argument is that this is effectively a marketing event for OpenAI — Tao chose to share via a ChatGPT link (using OpenAI's sharing infrastructure) rather than posting the transcript to his own blog or as a text document. The promotional effect for the product is real, and it's worth noting that the framing is partly shaped by OpenAI's product decisions. That doesn't invalidate the workflow, but it's a reasonable thing to hold in mind.
On the other side, several practitioners described similar workflows they've been running quietly. One commenter mentioned using Claude to stress-test mathematical proofs before journal submission. Others described using AI-generated adversarial questioning to strengthen technical design documents before peer review. The pattern across these accounts is consistent: expert leads, AI interrogates, human verifies.
The most analytically interesting thread was a debate about whether the model is "really reasoning" or doing something more like structured text production that mimics reasoning. The community divided roughly along the lines of the usual LLM epistemology debate. Our read: for the question of whether this is a productive tool, the mechanism matters less than the outcome. If running these sessions produces better-examined work, that's the relevant result. The philosophical question about whether the model understands mathematics is genuinely interesting and separately worth engaging, but it's not what determines whether the workflow is useful.
Several commenters also flagged that Tao has been publicly interested in AI for mathematics for some time — he's discussed it in blog posts and interviews, and this shared conversation is consistent with a documented interest rather than a sudden conversion.
Risks and Things to Watch
The risks here are real. Taking "Tao used it" as license to trust AI output uncritically in expert domains would be a significant error.
Hallucination in expert contexts remains dangerous, and the danger is specifically that it's hard to catch. Fluently incorrect algebra looks like correct algebra on the page. A consultant who isn't expert enough in the domain they're analyzing may not catch the places where the model's reasoning has quietly slipped. The workflow only works when the human can verify the output — and that requirement doesn't disappear because the tool is sophisticated.
The "just ask it" trap is the failure mode we're most concerned about for small teams. The workflow described here is disciplined and expert-led. The distorted version — "AI can handle the hard parts, I'll review the output" — flips the dynamic in a way that produces confident-sounding wrong analysis. The human needs to be driving throughout, not reviewing at the end.
Session drift is a documented behavior of current models that practitioners often underestimate. Over very long conversations, models can inconsistently apply constraints established early in the session. A logical premise set at message 3 may be violated at message 28 without the model signaling the inconsistency. The mitigation — periodic explicit restatement of the working framework — is necessary for long analytical sessions.
Data sensitivity is an operational concern that doesn't get enough attention in the excitement around these capabilities. Tao shared a mathematics conversation; the sensitivity is low. Teams using this workflow for competitive strategy, client analysis, or legal reasoning are sending that information to external servers. Enterprise tiers with explicit data-not-used-for-training commitments exist at higher cost. Free and standard tiers should be treated with appropriate caution for sensitive work.
Finally, there is real risk in the broader hype cycle around AI and mathematics specifically. Verified results in this space go through peer review. Viral moments do not. Treat any specific claim about AI-assisted mathematical breakthroughs on its specific merits, with attention to whether it has been independently reviewed. The Jacobian Conjecture has been "solved" before, multiple times, by people who weren't using AI and had excellent credentials. Peer review exists for a reason.
Frequently Asked Questions
Is this an actual mathematical breakthrough? Did Tao find a counterexample to the Jacobian Conjecture?
Based on what's available from the shared conversation and context, Tao appears to be working through and investigating a specific candidate structure that might serve as a counterexample — not announcing a verified and peer-reviewed result. The Jacobian Conjecture has a long history of near-misses and claimed results that didn't survive scrutiny. A genuine verified counterexample would represent major mathematical news and would need to go through the peer review process before the mathematics community considered it established. The ChatGPT session looks like part of an exploratory process. Treat any strong claim about the conjecture's resolution with significant skepticism until it's been reviewed by independent experts in the field.
Can this workflow be useful for people who aren't experts in their domain?
The workflow scales with expertise, and that's not an accident. What makes it productive is the practitioner's ability to catch the model's errors and redirect when the reasoning goes wrong. Someone without domain expertise running the same session is in a harder position — they can't reliably catch errors that the model embeds in otherwise reasonable-looking output. The workflow is still useful for developing expertise, but external verification of the model's outputs becomes more important, not less. You can't fully trust your own ability to catch errors in territory you're still learning.
What makes this different from just searching for information about a topic?
The structural difference is interactivity and specificity to your situation. A search engine retrieves documents that already exist. An AI reasoning session generates responses structured around the specific problem you've set up — including premises and constraints that are unique to your situation and that no existing document addresses. For problems that are genuinely novel to you, or decisions that don't have established documented answers, that difference is substantial. The model isn't retrieving; it's generating in context.
Should I use ChatGPT specifically, or would another model work?
For the reasoning-partner workflow, the relevant models are ChatGPT's o-series reasoning variants and Claude Sonnet/Opus. Both handle structured analytical dialogue well enough for this purpose. ChatGPT's reasoning models tend to be stronger on formal logical structure and mathematical consistency; Claude tends to be stronger on maintaining subtle contextual constraints over very long sessions. Try both on a real problem from your domain and see which reasoning style fits how you think. Many serious practitioners use both.
Is there a risk that this creates false confidence in AI-generated analysis?
Yes — and this is the most important risk to flag. The Tao workflow works precisely because Tao is skeptical of the output and verifies it against independent reasoning. If the main effect of this story is that people start trusting AI analytical output more uncritically, the net effect could be negative. The session structure requires the human to be the final arbiter of every claim made in the conversation. AI output is provisional material for expert review; it's not a conclusion.
How do I structure one of these sessions practically?
Start by stating your domain, the problem, and the current state of your thinking — all in the first message, before any question. Then proceed as an iterative dialogue, not a Q&A session. Explicitly ask the model to challenge your reasoning rather than confirm it. Periodically ask it to restate the logical framework you've established together, to check for drift. End by asking it to list the most important unverified claims in your reasoning — those become your external research action items.
What does this mean for the "AI replacing expert work" narrative?
Tao's session is actually evidence against the strong replacement thesis. He used a state-of-the-art model on a problem he has substantial independent expertise on, in one of the hardest intellectual domains that exists. The model didn't formulate the mathematical question, construct the candidate counterexample, or decide when the reasoning was sound. It provided structured response. Genuinely valuable — but the value is only unlocked when the expert is in charge. The more accurate frame isn't "AI replaces expertise" but "AI makes expertise more productive, provided the expert knows how to use it." That's a meaningful distinction, and this story illustrates it well.
Is the shared conversation link actually available, and can I read the full exchange?
The conversation was shared via OpenAI's public link-sharing feature, which means it should be accessible at the original URL as long as OpenAI keeps the link active. Content in shared ChatGPT links can sometimes be removed or expire. If the link goes down, the content may be preserved in archives or discussed in detail by the HN community. The broad contours of the session are well enough described in the HN thread to understand the workflow pattern even without reading the original.
Final Verdict
For small teams, freelancers, and agencies, the Tao story is best read as a strong practical signal about how to use AI — not as a wonder-of-technology moment, and definitely not as a reason to trust AI outputs more.
What it confirms, meaningfully, is that using AI as a structured reasoning partner is legitimate professional work. Not for content generation. Not for draft emails. For hard analytical problems where the quality of your thinking directly determines the quality of your outcomes. If you've been confining AI use to summarization and first drafts, this story is a legitimate prompt to reconsider whether you're leaving substantial value on the table.
Who should act on this now: anyone whose work bottleneck is reasoning quality rather than information access. Strategy consultants, product managers choosing between competing roadmap directions, developers making architectural decisions, analysts trying to find the flaw in their own frameworks before a client does. For these people, the reasoning-partner workflow is worth adopting immediately. The upfront investment is one disciplined session where you practice premise-setting and adversarial interrogation. Most people find that after two or three sessions, the pattern is intuitive and the return on session time is clear.
Who should be more cautious: teams where AI use is primarily for factual accuracy in external-facing content, or where the domain is highly regulated. This workflow doesn't reduce hallucination on factual claims — it uses the model for something where verified factual accuracy in the traditional sense isn't the primary output. Know what problem you're solving.
The broader signal worth sitting with: when Tao uses a tool explicitly and publicly, it doesn't mean you should trust the tool's outputs the way he can. What it means is that the workflow is worth understanding on its own terms. The specific method — expert judgment leading, AI providing structured response, human verifying everything — is transferable. What the Jacobian Conjecture session actually illustrates is something counterintuitive: the more expert you are in a domain, the more useful these tools become, because expertise is what lets you correct the model when it drifts. That inversion of the usual "AI handles what experts don't want to do" narrative is the most important analytical takeaway here. AI makes expertise more powerful. It doesn't substitute for it.