↓Skip to main content
  1. Posts/

Did OpenAI Steal Mathematicians’ Work? r/LocalLLaMA Has Thoughts.

·4 mins

The Allegation: OpenAI vs. Mathematicians #

It’s the kind of drama we all live for in AI circles. Earlier this week, a post blew up on r/LocalLLaMA alleging that OpenAI might’ve scraped work from mathematicians without permission—or even attribution. The accusation? OpenAI trained its models, like GPT-4, on datasets containing academic papers, including niche, cutting-edge mathematical research, but snubbed the people creating that work.

One user, u/primefact_scholar, pointed out that “it’s not just papers, it’s datasets and proofs embedded in proprietary text. These tools are regurgitating results it couldn’t possibly derive on its own.” Strong accusation, but let’s not pretend this is unheard of in AI training practices.

Scraping Academic Knowledge: Where’s the Line? #

We all know large AI models munch on absurdly huge text datasets to get smart. That includes everything from Reddit debates to obscure 90s fanfiction on LiveJournal. But mathematical research? It stings.

Here’s why:

  1. Math isn’t fluff. It’s not like copying pop culture blog posts. This stuff represents years (sometimes decades) of labor.
  2. Accuracy matters. These models don’t always spit out correct solutions, but people use them anyway. Now mathematicians risk being misquoted or—worse—having their findings diluted with pseudo-math.

u/LaTeXLand wrote: “The problem is, once GPT does a bad proof, people assume the fault lies with the human paper it mirrors. It’s reputational damage.”

When asked about sources, OpenAI hasn’t been super clear. They claim they filter for publicly available data, but what counts as “public”? ArXiv preprints? Lecture notes someone forgot were online? You can smell the gray area here.

The “Non-Attribution Problem” #

Another bone of contention? The lack of citations. A few voices in the thread stated that if GPT-4 borrows heavily from an academic’s unique terminology, or replicates long proofs, shouldn’t it cite them? That’s the bare-minimum etiquette academics expect.

u/mathnerd42 shared an example where GPT-4 generated a sophisticated proof for a combinatorial graph problem, “eerily similar to a paper I co-authored in 2021.” No citation. No acknowledgment. Just vibes.

Sure, OpenAI doesn’t claim the outputs as fact—but neither do users. And trust me, lazy grad students outnumber careful researchers.

Why This Matters to YOU (Yes, You!) #

Okay, maybe you’re thinking, “I’ll never need AI to crank out topology proofs; what’s the big deal?” Here’s why it matters:

When AI arbitrarily scrapes everything, those datasets reflect flaws, biases, and permission violations. This issue doesn’t stop with math. It’s precedent. If OpenAI gets away with sweeping academic work under the rug, what’s stopping them from eating medical research burritos next, claiming your groundbreaking trial as their training fuel?

Legal and ethical frameworks for dataset curation lag far behind the tech. u/harmoniclawyer commented: “Do we need a DMCA for research papers? ‘Cause that’s where this is heading.”

What Can Researchers Do? #

The community had some practical suggestions—but they’re all tough pills to swallow.

  1. Lockdown access. Many suggested moving sensitive or novel work behind stricter paywalls. ArXiv would hate this, but journalists, web scrapers, and GPT models hate it more.
  2. Watermarking ideas. Embed unique or random patterns into publications—a sort of intellectual fingerprint. This feels dystopian, but it’s gaining traction.
  3. Public ethics pressure. The loudest voices on r/LocalLLaMA argued OpenAI should face more public scrutiny. Can transparency be demanded? Or is this like yelling at clouds?

FAQ #

Can GPT-4 really output math proofs? #

Yes, but often incorrectly, especially for anything too niche. It’s shockingly good at looking convincing though. Always verify.

Is this stealing, legally speaking? #

Depends. OpenAI scrapes anything “publicly available,” meaning if you can click it without a login or decryption key, it’s fair game under current laws. Ethically? Way murkier.

How do I stop OpenAI from using my work? #

You can’t, realistically. Website no-crawl headers (robots.txt) might help, but they don’t guarantee LLMs won’t grab your data. Write to OpenAI or join academic consortiums pushing for dataset regulations.


Final thought: OpenAI needs to get its act together. Math nerds, keep being mad—this isn’t a small issue.