Where your AI's training data comes from — and why it matters for your reputation
AI tools were trained on content scraped from the public web — raising real questions about copyright, output ownership, and commercial risk for Australian SMEs.
If you've been using an AI tool to write marketing copy, summarise contracts, or draft client reports, you probably haven't stopped to ask where the model learned to do that. Most people don't. The tool works, the output looks reasonable, you save a bit of time, and that's the end of the mental loop.
The honest version is a little more complicated. Every large language model was trained on a dataset — usually a very large one, assembled from the public internet, books, code repositories, and other sources. The decisions made during that data collection process have downstream consequences that can affect your business: legally, commercially, and reputationally. Not dramatically, and not in ways that should send you running from the tools. But in ways worth understanding before you put AI-generated content in front of clients or customers.
This post covers where training data actually comes from, what the copyright situation looks like from an Australian-business perspective, who owns what you generate, and where the genuine commercial risks sit — without the doomsaying or the breathlessness.
Where does AI training data actually come from?
The bulk of training data for most large language models comes from large-scale web scrapes — essentially automated programs that crawl publicly accessible websites and collect text at scale. Common Crawl, a non-profit dataset of web content built up over more than a decade, forms the backbone of many models' training sets. Beyond that, datasets typically include digitised books, Wikipedia, code from public repositories, Reddit threads, academic papers, and news articles.
The scraping happens at a scale that's genuinely difficult to conceptualise. We're talking about an enormous volume of text, assembled from content that was mostly published with no expectation that it would be fed into a model. The websites involved range from major publishers to small business blogs to individual writers who've been doing their craft online for twenty years.
Most AI labs have published high-level descriptions of their training data sources — names like Common Crawl, The Pile, or their own named datasets — but granular disclosure of exactly what content was included is rare. That opacity is itself a source of ongoing legal and ethical tension, and it's why a growing number of creators, publishers, and news organisations globally have launched litigation or negotiated opt-out or licensing arrangements with the major labs.
The practical upshot for a small business owner: the model you're using almost certainly learned from content that was scraped without the creators' explicit consent. Whether that constitutes copyright infringement — or fair use, or something else entirely — is a live and unresolved legal question in multiple jurisdictions, including Australia.
Does Australian copyright law cover this, and does it affect you?
Australian copyright law is broadly similar to other common-law jurisdictions, but there is no explicit AI training data exemption in the Copyright Act 1968. This matters because it means the legality of scraping and training on publicly available content is genuinely unsettled — there has been no definitive Australian court ruling on whether training a model on scraped content constitutes infringement.
The Australian government has been consulting on AI and copyright through the Attorney-General's Department, and the Australian Copyright Council has published guidance flagging the uncertainty. In October 2025 the federal government ruled out a text-and-data-mining exception for AI training, so there's no special carve-out coming; how existing law applies to models trained overseas is still untested in court. The short version: it is not clearly legal, it is not clearly illegal, and the law is still catching up to the technology. The UK is in a similar spot. The EU took a different route: under a 2019 directive, it allows text and data mining unless the rights holder opts out. The US has a developing body of case law, but nothing conclusive at scale yet.
Where this becomes directly relevant to your business is not in the training process itself — you didn't build the model — but in what you do with the output. If a model was trained on copyrighted creative work and can reproduce portions of that work closely, then content you publish that incorporates those portions could, in theory, expose you to claims. This is an edge case for most business uses — writing a supplier email or summarising a meeting transcript isn't going to reproduce anyone's novel — but it is a genuine risk for content that closely mimics a distinctive style, or for images generated in the style of a living artist.
The practical advice here is straightforward: treat AI-generated content the way you'd treat content from a contractor you can't fully vet. Review it, edit it, and don't publish it verbatim without reading it. That's good practice for quality reasons anyway.
Who owns the content that an AI generates for you?
This is the question most small business owners care about most, and the honest answer is: it depends on the tool, and the law hasn't fully caught up. In Australia, the Copyright Act 1968 does not currently recognise AI as an author. Copyright generally requires a human author, which means purely AI-generated content — where there was no meaningful human creative contribution — may not attract copyright protection at all. You'd own something that anyone could legally copy.
In practice, most business outputs involve enough human direction, editing, and selection that a reasonable argument for copyright can be made. If you write a detailed brief, refine the output through several rounds, and edit the final product substantially, you're probably the effective author in any practical sense. If you hit "generate" and paste the result directly into your website, the position is murkier.
The terms of service for the major AI tools also matter here. Most major providers — and this is worth actually reading — grant you broad rights to use the output commercially, but some include clauses around indemnification, third-party claims, and what happens if the output turns out to infringe on something. These terms vary between tools and change over time. If you're using AI-generated content at any scale in a commercial context, a thirty-minute read of the relevant terms of service is worth your time.
The safest operating position for an Australian SME is to treat AI output as a strong first draft that belongs to you once you've shaped it, rather than as a finished product that fell out of the machine.
What are the actual commercial risks — and how real are they for a small business?
Let's be specific about risk, because the internet tends toward the extremes — either "AI is totally fine, don't worry" or "you'll be sued into oblivion." The realistic picture for an Australian SME sits somewhere much less dramatic.
1. Reputational risk from recognisable content. If an AI tool reproduces a distinctive phrase, passage, or creative element from a known source and you publish it under your name, you look like you've copied someone's work — even if you didn't know. This is unlikely for most business content, but it's a real hazard for marketing copy, blog posts, and anything written in a distinctive creative voice.
2. Competitive risk from generic output. This is more common and more immediately damaging for most businesses. If you and every competitor are feeding the same prompt into the same model, you'll get similar-sounding copy, similar-looking structure, and similar ideas. Your content stops being yours in any meaningful sense. AI helps with volume; it does not automatically help with distinctiveness.
3. Disclosure expectations are shifting. Some industries — journalism, academic publishing, some government procurement — are already requiring disclosure of AI use. Legal and financial services are watching this closely. If your sector moves toward mandatory disclosure and you haven't thought about it, you'll be caught flat-footed. It's worth tracking what your industry body is saying now, before it becomes a compliance requirement.
4. Supplier and client contract clauses. An increasing number of enterprise clients and government buyers are including AI-use clauses in contracts — sometimes prohibiting it for certain deliverables, sometimes requiring disclosure. If you're delivering content, reports, or documents under a contract and using AI to produce them, check whether your agreement covers it.
How should this change how you actually use AI tools?
None of this is a reason to stop using AI tools. It is a reason to use them with a bit more intentionality than most people currently apply.
The most practical shift is to treat AI output as raw material, not finished product. That means a human — you, a team member, someone with genuine domain knowledge — reads, edits, and takes responsibility for everything that goes out under your name. Not as a legal formality, but because the output is genuinely better when someone who knows the subject has touched it. The model doesn't know your clients, your market, or your voice. You do.
For creative or marketing content specifically, the risk of generic sameness is more immediate than any copyright concern. The fix isn't complicated — use AI to get from zero to a rough shape, then bring your own perspective, your own examples, and your own observations. That combination is where the output stops being forgettable.
For commercial agreements — contracts, proposals, anything with a client's name on it — run a sensible review process regardless of whether AI touched the first draft. That's just professionalism, and it protects you from the occasional AI-generated hallucination that reads plausibly but is quietly wrong. If you're using AI automation for document-heavy workflows, the human review step should be built into the process design, not bolted on as an afterthought.
On data residency and privacy: training data questions are distinct from what happens to the information you put into a model when you use it. If you're putting client details, financial records, or anything sensitive into an AI tool, the question of where that data goes is separate — and arguably more urgent for most small businesses. That question is covered in more detail in where your business data goes when you use AI.
What's the practical summary for an Australian SME owner?
The training data situation is genuinely unsettled — legally, ethically, and commercially. The laws haven't caught up to the technology, the major labs have been opaque about what they scraped, and the copyright question for AI-generated output remains unresolved in Australia. None of that means AI tools are unusable. It means they come with caveats worth knowing.
The caveats that actually matter for a small business are: review everything before publishing, don't rely on AI-generated content as automatically yours in a copyright sense, watch for generic sameness in anything customer-facing, and read the terms of service for whatever tool you're paying for. The litigation happening between publishers and AI labs in the US and UK will eventually produce clearer rules — in Australia, that process is slower, but it's moving.
In the meantime, the operating principle is the same one that applies to most business decisions: know what you're buying, know what you're signing, and don't outsource your judgement to the tool. AI is useful precisely because it handles the mechanical parts of a task quickly. The parts that require your name on the output — and your reputation behind it — still need you in the loop.
Common questions
Was the AI tool I'm using trained on copyrighted content?
Almost certainly, yes. Most large language models were trained on content scraped from the public web, including copyrighted articles, books, and creative work, typically without the creators' explicit consent. Whether this constitutes infringement under Australian law is unresolved — the Copyright Act 1968 has no explicit AI training data exemption.
Who owns content that an AI generates for my business?
Under the Copyright Act 1968, Australian copyright generally requires a human author, so purely AI-generated content may not attract copyright protection. If you direct, shape, and substantially edit the output, you have a stronger claim. Check the terms of service for your specific tool — they vary significantly on commercial use rights.
Can I get in legal trouble for using AI-generated content in my business?
For most typical business uses — emails, summaries, internal documents — the practical risk is low. The higher risks are: AI reproducing recognisable content from a known source, publishing generic copy that's indistinguishable from competitors, and contract clauses with clients or government buyers that restrict or require disclosure of AI use.
Does the Privacy Act 1988 apply to AI tools I use for my business?
The Privacy Act 1988 applies to businesses with over $3M turnover (with some exceptions). If it applies to you, how you handle personal information — including what you enter into AI tools — matters. But training data questions are separate from what happens to data you input into a tool during normal use. See where your business data goes when you use AI for the data-residency angle.
Do I need to disclose to clients that I used AI to produce content or documents?
There is no blanket Australian legal requirement to disclose AI use, but some industries and enterprise clients are beginning to include AI-use clauses in contracts. Journalism, academic publishing, and some government procurement already have disclosure expectations. Check your contracts and watch your industry body's guidance — requirements are shifting.
See if Neurastruct can help your business
Book a free 30-minute consultation
No commitment. We'll walk through your biggest admin time-sucks and whether AI is the right fit for your specific business.
Book a consultation
Peter McLean
Founder, Neurastruct
Australian small-business operator since 2001 and 16 years as a national account manager; AI certificates from Anthropic (2026) and Google (2025).