Glossary · L

LLM (Large Language Model)

The underlying technology behind most modern AI tools — a model trained on vast amounts of text to predict and generate language — and the reason nearly every AI risk in this glossary traces back to the same basic fact: an LLM processes whatever text it's given, sensitive or not, without knowing the difference on its own.

What Is an LLM?

A large language model (LLM) is a type of AI model trained on vast amounts of text data to understand and generate human language, powering the chat assistants, drafting tools, summarization features, and AI agents that have become common across business software. An LLM works by predicting the most statistically plausible continuation of a given text based on patterns learned during training, which is what allows it to answer questions, summarize documents, draft emails, and hold conversations — but it's also the underlying mechanism behind both of the model's most consequential limitations: it doesn't inherently distinguish sensitive data from ordinary data, and it doesn't retrieve verified facts so much as generate plausible-sounding text, which is the root cause of hallucination.

Understanding LLMs at this basic level matters for AI risk specifically because most of the risk vectors, controls, and compliance obligations covered elsewhere in this glossary exist precisely because of how LLMs work. An LLM given a prompt containing a customer's name and account number will process that information exactly the same way it processes any other text — it has no built-in concept of "this is sensitive, handle it differently" unless a system around it is specifically built to detect and protect that content before it reaches the model.

Practical Industrial Use

A company deploying a customer service chatbot built on an LLM is a clear example of both the model's value and its blind spots operating simultaneously. The LLM is what makes the chatbot capable of understanding a customer's natural-language question and generating a coherent, helpful response — a genuine and significant capability. But the same LLM, if a customer's message includes their account number or personal details as part of their question, processes that sensitive information with no more built-in caution than it applies to any other part of the conversation, because the model itself doesn't have a mechanism for recognizing and specially handling sensitive content unless something upstream does that work first.

The same dynamic underlies every AI tool discussed elsewhere in this glossary: an AI scribe's LLM processes a patient's health information the same way it processes any dictated text; an AI claims tool's LLM processes a customer's financial details with the same mechanism it uses for any other document; an AI legal tool's LLM generates a case citation with the same fluent confidence whether the case is real or entirely fabricated. The risk vectors, anonymization needs, and human-oversight requirements described throughout AI governance all trace back to this same underlying property of how LLMs actually function.

What Happens Without It

Organizations that don't understand the basic behavior of the LLM underlying their AI tools tend to build governance programs around the wrong assumptions — treating an AI tool as if it has some inherent judgment about what's sensitive or what's factually verified, when the model itself has neither of those capabilities built in in the way a human employee's judgment would. This misunderstanding is what leads directly to the specific risk vectors covered throughout AI governance: sensitive data reaching a model unprotected because no one built a safeguard for a model that was never going to provide one itself, or hallucinated content reaching a real decision because no one verified it, on the assumption that the model's fluent, confident output meant it was also accurate.

⚠ Risk Without Protecting LLM Interactions This gap compounds because LLM-powered tools are often adopted for their conversational fluency and apparent understanding, which can create a false impression that the model is reasoning and judging content the way a person would, rather than generating statistically plausible text based on patterns in its training data. An organization that anthropomorphizes an LLM's capabilities in this way is more likely to skip the specific external safeguards — anonymization, human review, audit trails — that the model's actual underlying mechanism makes necessary.

With LLM Behavior Understood and Governed

  • Sensitive data is anonymized before reaching an LLM, since the model itself has no built-in way to recognize and protect it
  • AI-generated output involving factual claims is independently verified, since the LLM generates plausible text rather than retrieving guaranteed facts
  • Governance programs are built around what LLMs actually do, not an assumption of human-like judgment they don't possess
  • The same underlying risk — an LLM processing whatever it's given — is addressed consistently across every AI tool built on one

Without It

  • Sensitive data reaches LLM-powered tools unprotected because the model was mistakenly assumed to handle it appropriately on its own
  • Hallucinated output is trusted at face value because the model's fluent tone is mistaken for verified accuracy
  • Governance programs miss the actual mechanism creating AI risk, addressing symptoms rather than the underlying cause
  • Every new LLM-powered tool inherits the same unaddressed risk, since the root behavior driving it was never actually understood

How This Relates to Questa AI

Questa AI is built specifically around the reality of how LLMs actually work — that they process whatever text reaches them without inherently distinguishing sensitive content, and that they generate plausible text rather than guaranteed facts. Its entity-detection engine anonymizes sensitive data before it reaches any LLM-powered tool, providing the external safeguard the model itself doesn't provide, regardless of which specific LLM or vendor is powering a given AI feature.

This underlying understanding is also why Questa's Safe AI Agent controls emphasize human oversight rather than assuming an LLM's output can be trusted without verification — the goal is to build governance around the model's actual behavior, not an idealized version of what an AI tool might be assumed to do. Combined with Blackbox's documented record of what an LLM actually processed and generated, and the governance dashboard's visibility across every connected AI tool, Questa treats the underlying mechanics of LLMs as the starting point for its governance approach, rather than treating each AI tool's risk as unique and unrelated to this shared root cause.

Frequently asked questions

No, not inherently. An LLM processes text based on patterns learned during training and doesn't have a built-in mechanism to recognize that a given piece of text is a customer's account number or health information versus any other content, unless a separate system is specifically built to detect and protect that data before it reaches the model.

It's a general property of how LLMs generate text — predicting plausible continuations rather than retrieving verified facts from a database — which means hallucination shows up across LLMs from different vendors and of different capability levels, though more advanced models generally hallucinate less often on many tasks.

No. More advanced LLMs may perform certain tasks more capably, but the fundamental behavior — processing whatever text it's given without inherent sensitivity detection, and generating plausible rather than guaranteed-accurate text — remains true regardless of the model's overall capability level.

The LLM is the underlying model doing the language processing; the AI tool or product is the broader application built around it, which may include additional features like retrieval from a specific database, safety filters, or, ideally, external safeguards like anonymization and human oversight that address the LLM's inherent limitations.

Because governance controls that don't account for how LLMs actually process and generate text tend to address the wrong problem — assuming the model has judgment or verification capabilities it doesn't have, rather than building the external safeguards the model's actual behavior requires.

Many of the most common AI risk vectors — sensitive data exposure and hallucination in particular — trace directly back to fundamental LLM behavior, though some risks, like over-permissioned AI agents or third-party vendor data retention, relate more to how an AI system is deployed and governed around the model rather than to the LLM's core language-processing mechanism itself.

See LLM (Large Language Model) in practice

Questa AI anonymizes sensitive data before it reaches any AI model — across documents and live prompts, with governance and data-residency control.

Contact

Contact Us

Have questions or ready to explore how Questa AI can transform your business?