The technical vocabulary — what the systems actually are and how they are built, in plain language.
Writing and structuring a page so a machine can lift a single clear answer out of it and say that answer back — in a featured snippet, a voice assistant reply, or an AI summary at the top of a search page. The win is being the answer, not ranking in a list of links.
OriginNo single documented coiner
Credited to Jason Barnard of Kalicube, who says he coined it in 2017 and set the method out in a January 2018 white paper, “The New Face of SEO: Answer Engine Optimization,” written with Chee Lo at Trustpilot. Treat this as a claim by the person making it: the coining is his own account, secondary write-ups date it to 2017 or 2018 depending on who you read, and Google itself does not use the term. Marked uncertain for that reason.
Primary source
The repeating cycle an AI agent runs: read the situation, generate an action (such as code), run it, check the result, and go again. Most AI coding tools are built around this loop.
OriginNo single documented coiner
A descriptive phrase from AI-agent engineering with no single coiner.
Everything wrapped around an AI model to turn it into something that can actually do a job: the loop that lets it take one step after another, the tools it is allowed to call, the memory of what it has already done, the sandbox it runs inside, and the limits and approvals that stop it running away with your money. The usual shorthand is “Agent = Model + Harness.”
OriginNo single documented coiner
No single person coined “harness” in this sense — it was inherited from the older software “test harness” and from machine-learning “evaluation harnesses” such as the one used to score SWE-bench, then stretched to cover the whole wrapper around a model. The UK’s AI Security Institute was already describing an agent as a model plus its scaffolding in 2023. The name for the discipline, “harness engineering,” is credited to Viv Trivedy, whose “Anatomy of an Agent Harness” post Addy Osmani points to as the clearest derivation. Usage is still loose: people call Claude Code, Codex CLI and an evaluation script all “harnesses,” which is why a 2026 arXiv paper set out to define the boundary.
Primary source
Explained in plain words on AI Basics
AI systems designed to act semi-autonomously: setting or interpreting goals, planning a sequence of actions, and using tools to complete multi-step tasks, rather than just answering a single question.
OriginNo single documented coiner
“Agent” is long-standing academic AI terminology; the current “agentic AI” usage grew out of industry practice from 2023 onward rather than a single coining document.
Explained in plain words on AI Basics
Key resources in the library
A hypothetical AI system that could understand, learn, and perform any intellectual task a human can, rather than excelling at narrow, specific tasks.
OriginNo single documented coiner
Earliest documented use is Mark Gubrud’s 1997 paper on nanotechnology and security; the term was independently reinvented and popularized around 2002 by Shane Legg (later a DeepMind co-founder) and Ben Goertzel, unaware of Gubrud’s earlier use — Legg has said “I didn’t invent the term, I reinvented it.” Both attributions are commonly cited depending on whether “coined” or “popularized” is meant.
Deliberately attacking an AI system to find ways it can be made to fail, misbehave or cause harm — for example, coaxing a chatbot into giving dangerous instructions, leaking data, or ignoring its rules — so the problems can be fixed before release. Red teams may be internal staff, outside experts or members of the public at organized events.
OriginNo single documented coiner
Borrowed from military and later cybersecurity practice, where a “red team” plays the adversary against the defending “blue team.” AI labs adopted the term for pre-release testing around 2020–2022, and the 2023 U.S. executive order on AI and the DEF CON Generative Red Team event helped make it standard vocabulary.
Primary source
A fixed set of step-by-step instructions a computer follows — a recipe. In headlines, 'the algorithm' usually means the ranking system that decides what you see on social media, which is really many algorithms plus a learned model.
Origin
Named after the 9th-century Persian mathematician Muhammad ibn Musa al-Khwarizmi; 'al-Khwarizmi' became 'algorismus' in medieval Latin, then 'algorithm'. The modern computing meaning was settled by the mid-20th century.
The part of an AI product people actually touch — the app, chatbot or feature — built on top of someone else's AI model. A company like a legal-document assistant sits in the application layer; the model underneath (GPT, Claude and so on) is the 'foundation layer'.
OriginNo single documented coiner
Borrowed from older networking and software terminology, where the 'application layer' is the top of the stack. In AI it spread through investor and startup writing from about 2023 as a way to sort companies into model-makers versus app-builders. No single documented coiner for the AI usage.
The overall design of an AI model or system: how its parts are arranged and connected. 'Transformer architecture' means the specific design behind GPT, Claude and most modern AI; a new architecture means a different design, not just a bigger version.
OriginNo single documented coiner
Borrowed from building design via early computing — 'computer architecture' was standard by the 1960s (IBM's System/360 era). Applied to neural networks from the 1980s onward. No single documented coiner for the AI usage.
AutoRegressive Integrated Moving Average — a classic statistics method for forecasting a number over time (sales, prices, demand) from its own past values. It is not modern AI, but it is still a common baseline that AI forecasting tools are measured against.
Origin
Developed by statisticians George Box and Gwilym Jenkins in their 1970 book 'Time Series Analysis: Forecasting and Control'.
The science and engineering of building machines that can perform tasks normally requiring human intelligence: reasoning, learning, problem-solving.
An open, federated protocol for social networking — the technical foundation of Bluesky. Instead of one company owning the network, many servers (PDSes) host user data, and relays aggregate it into a shared stream anyone can read or build on. User accounts are portable: you can move between providers without losing your posts, follows or social graph.
Origin
Created by Bluesky Social PBC, originally as “ADX” in spring 2022 and renamed the Authenticated Transfer Protocol in October 2022. The initial blog post by the Bluesky team described account portability, algorithmic choice and interoperation as its goals. An earlier name, “Bluesky,” came from a Twitter (later X) initiative announced in 2019; Bluesky became an independent company in 2022.
Primary source
Software employers use to collect, sort and manage job applications. An ATS stores résumés, moves candidates through hiring stages and often ranks or filters applicants. Newer systems add AI features that compare a résumé's meaning to the job description instead of only counting exact keywords.
OriginNo single documented coiner
An HR-software term that spread in the 1990s and 2000s as companies moved job applications online. No single documented coiner.
Primary source
A technique that lets a model dynamically focus on the most relevant parts of its input when producing each part of its output.
An AI model that generates output one piece at a time, where each new piece is predicted based on everything generated so far. This is how most text-generating AI actually writes: predicting the next word based on all previous words, one word at a time.
Origin
“Autoregression” is a statistics term predating AI by nearly a century, introduced by statistician Udny Yule in a 1927 paper modeling a value as a function of its own past values. The term was carried into machine learning as sequence models were described as generating output step-by-step from prior output.
Primary source
Explained in plain words on AI Basics
Bringing data into a system in scheduled chunks — for example, loading yesterday's sales every night at 2 a.m. Simpler and cheaper than streaming, but the data is always somewhat out of date.
OriginNo single documented coiner
Descends from batch processing on early mainframes (1950s–60s), where jobs were queued and run in groups. 'Ingestion' as a data-pipeline word is later industry jargon with no single source.
A program that translates code written in the C language into instructions a computer's processor can run. Building one is a classic hard test of programming skill, which is why AI labs now use it to show off what coding agents can do.
Origin
Dennis Ritchie created C and its first compiler at Bell Labs in 1972–73. The word 'compiler' is credited to Grace Hopper, who used it in the early 1950s.
Primary source
A general-purpose programming language that extends C with object-oriented features, templates and direct memory control. It is used where raw speed matters: operating systems, game engines, browsers and the performance-critical layers of AI frameworks such as PyTorch and TensorFlow, which are written in C++ under their Python interfaces.
Origin
Created by Bjarne Stroustrup at Bell Labs, who began work in 1979 on “C with Classes” and released the first commercial version in 1985. The name “C++” (the increment operator in C) was suggested by Rick Mascitti in 1983.
Primary source
Everyday shorthand for an AI model using function calling — reaching out to a search engine, calculator, code runner or other app to get something done. Also called 'tool use' or 'tool calling'. A model that can call tools is the basic ingredient of an AI agent.
OriginNo single documented coiner
Developer slang that grew out of 'function calling' and 'tool use' in 2023–2024; no single coiner.
In AI tools, a workspace beside or instead of the chat where the AI and the person edit the same document, code file or design directly, rather than passing messages back and forth. OpenAI, Anthropic (as Artifacts), Google and GitHub Copilot offer versions of it.
OriginNo single documented coiner
Borrowed from the painter’s canvas via earlier software usage (such as the HTML canvas element). OpenAI launched a feature named Canvas in ChatGPT in October 2024; GitHub describes canvases for Copilot on its blog.
Primary source
A prompting technique where a model is asked to write out its step-by-step reasoning before giving a final answer, which measurably improves accuracy on multi-step problems.
The three typed question formats in TypeSafe’s System One API. A Choice picks one option from a set you define and returns a probability for each. A Score rates content along ordered levels. A Noul is a yes/no question that returns the probability the answer is yes. Because the answer shape is fixed in advance, code can use it directly without parsing free text.
Origin
Defined in TypeSafe’s documentation. “Noul” is TypeSafe’s own coinage for its yes/no type; the documentation does not explain the name.
Primary source
Two meanings in AI. First, groups of similar data points that an algorithm finds on its own, without being told the categories — for example, sorting customers into look-alike groups ('clustering'). Second, a computer cluster: many machines, often thousands of GPUs, wired together to train or run large AI models as one system.
OriginNo single documented coiner
'Cluster analysis' was named in psychology by Robert Tryon in 1939, and the popular k-means method traces to Stuart Lloyd (1957) and James MacQueen, who coined 'k-means' in 1967. 'Computer cluster' grew out of 1960s–90s computing practice with no single coiner.
Meeting the rules and standards that apply to an AI system — a regulation's requirements, a company's own policy, or a contract's terms. In practice it means records, testing, disclosure, and the audits that prove them. A firm can 'govern' AI voluntarily and still be out of 'compliance' with a law.
OriginNo single documented coiner
A standard term in law, finance, and corporate management for adhering to rules and proving it. Applied to AI as regulations such as the EU AI Act added concrete obligations — risk assessment, documentation, human oversight — that firms must demonstrate. No single coiner.
Key resources in the library
- NIST AI Risk Management Framework
NIST, 2023 (updated)
The US National Institute of Standards and Technology's voluntary framework for managing AI risk across an organization: govern, map, measure, manage.
See it in the library - AI Best Practices for Authors — The Authors Guild
The Authors Guild, 2026
Guidance for authors on AI: disclosure, contract terms, copyright registration, and protecting work from unauthorized training use.
See it in the library - Publisher Policies on AI — Oklahoma State University Library
Oklahoma State University Library, ongoing
A library-maintained directory of how academic and trade publishers address AI in their policies, covering authorship, peer review, and training.
See it in the library
The processing capacity behind an AI system, treated as a quantity you can buy, ration or run out of: the chips, the data centers, the electricity and the networking that let a model be trained or answered on. It is used as a mass noun — “more compute,” “compute-constrained” — which is worth noticing, because it turns land, power and water into a single abstract number.
OriginNo single documented coiner
No coining event: engineers turned the verb into a noun, and the usage spread through machine-learning research and industry from the mid-2010s as training runs grew large enough that capacity became the limiting factor. It is now standard in government policy writing too, as in this Tony Blair Institute explainer, which defines it as the stack of hardware, software and infrastructure that stores, processes and transfers data at scale.
Primary source
The main United States federal anti-hacking law, passed in 1986. It makes it a crime to access a computer 'without authorization' or to exceed authorized access. Because the phrase 'exceeds authorized access' was long read broadly, the law has been invoked in cases ranging from data theft to violating a website's terms of service — including the prosecution of Aaron Swartz for downloading academic articles.
Origin
Enacted by the U.S. Congress in 1986 as an amendment to an earlier computer-crime law. In Van Buren v. United States (2021), the Supreme Court narrowed the law, ruling that 'exceeds authorized access' does not cover people who misuse information they are entitled to see.
Primary source
A piece of malicious software that copies itself from computer to computer across a network without anyone clicking anything. Unlike a virus, which needs a person to open an infected file, a worm spreads on its own. Worms matter to AI policy because the same self-spreading behavior is a worst-case scenario people worry about for AI systems that can write and run their own code.
Origin
The term comes from John Brunner's 1975 science-fiction novel The Shockwave Rider, which featured a self-propagating 'tapeworm' program. The first famous real worm was the Morris worm of 1988, built by graduate student Robert Tappan Morris, which accidentally crashed much of the early internet and led to the first conviction under the U.S. Computer Fraud and Abuse Act.
Primary source
The rules a database uses so that many people or programs can read and change data at the same time without corrupting it — for example, stopping two people from booking the last seat on a flight.
Origin
A core topic of database research from the 1970s. Eswaran, Gray, Lorie and Traiger's 1976 paper 'The Notions of Consistency and Predicate Locks in a Database System' set out two-phase locking, and Jim Gray's later work on transactions earned him the 1998 Turing Award.
Primary source
A rule that only lets an AI system act on its own when it is sure enough — for example, above a set confidence score. Below that line, the case is handed to a person or a slower check instead.
OriginNo single documented coiner
Borrowed from long-standing 'confidence threshold' practice in machine learning and quality control; the 'gate' wording is informal and has no single coiner.
A number, usually between 0 and 1, that a model reports alongside its answer to say how sure it is. It is useful for deciding whether to act automatically or send a case to a person. A confidence value is not a guarantee: calibration is measured across many predictions, so a confident single answer can still be wrong.
OriginNo single documented coiner
General machine-learning usage. TypeSafe’s documentation separates confidence (how concentrated the answer is) from the probability of each option.
Primary source
The amount of text (measured in tokens) a model can take into account at once during a single conversation or task.
OriginNo single documented coiner
Standard technical vocabulary that arrived with sequence models rather than a single coining event; no individual is credited.
Explained in plain words on AI Basics
A four-step structure for working with an AI chatbot as a thinking partner: give it Context about your situation, assign it a Role, let it Interview you with clarifying questions one at a time, and only then give it the Task. The interview step is the distinctive part — the AI asks the questions, which surfaces considerations the user would not have thought to include.
Origin
Introduced by Geoff Woods in his book The AI-Driven Leader (2024). Woods describes CRIT as Context, Role, Interview, Task in his own interviews and talks; some secondary summaries instead expand the acronym as challenging assumptions, reducing bias, improving strategy, and testing ideas, so the two expansions circulate side by side.
Primary source
Criticism that repeats or amplifies exaggerated technology claims while trying to challenge them. The critic rejects the promised future, but still treats the promoter's description of the technology as the starting point instead of checking what the system can actually do.
Origin
Coined by technology historian Lee Vinsel in his 2021 essay "You're Doing It Wrong: Notes on Criticism and Technology Hype." Vinsel used it for criticism that remains trapped inside the stories told by technology promoters.
Primary source
Nvidia's software platform that lets AI programs run calculations on its graphics chips (GPUs) instead of ordinary processors. Nearly all large AI models are trained on CUDA, which is a big reason Nvidia dominates AI hardware.
Origin
Created at Nvidia; Ian Buck and colleagues published the first public release in 2007, growing out of Buck's Stanford PhD work on the 'Brook' GPU programming language. CUDA originally stood for 'Compute Unified Device Architecture', though Nvidia later dropped the expansion.
Primary source
The study of how any system — a machine, an animal, a company, a country — steers itself using feedback: it acts, measures the result, and corrects. It is where ideas like control loops, self-regulation and “the system responds to its own output” come from, and it predates artificial intelligence as a field.
Origin
Named by Norbert Wiener in his 1948 book “Cybernetics: Or Control and Communication in the Animal and the Machine,” from the Greek kybernetes, a steersman. The ideas were worked out publicly at the Macy Conferences of the 1940s and 1950s alongside Margaret Mead, Heinz von Foerster and Warren McCulloch. AI largely set the word aside after the 1960s; this free 2024 IFAC paper argues that was a mistake, and traces the history.
Primary source
A subset of machine learning that uses large, multi-layered neural networks to automatically learn complex patterns from data, rather than relying on hand-designed features.
OriginNo single documented coiner
“Deep” network terminology developed across the neural-network research community; the phrase came into general use through the 2000s rather than from one paper.
The technique behind most modern image generators: starting from random noise and gradually “denoising” it into a coherent image, reversing a process trained by progressively adding noise to real images.
A sealed, portable package holding a program plus everything it needs to run, so it works the same on any machine. AI agents are often run inside containers so their mistakes can't damage the rest of the computer.
Origin
Docker was released by Solomon Hykes and the company dotCloud in 2013, building on older Linux container features.
Primary source
The human tendency to read real understanding, care or awareness into a system that is only matching patterns on the surface. It is why people disclose things to a chatbot they would not tell a colleague, and why a warm tone in an answer feels like evidence of a mind behind it.
OriginNo single documented coiner
Named after ELIZA, the 1964–66 script written by MIT computer scientist Joseph Weizenbaum, whose secretary reportedly asked him to leave the room so she could talk to it privately — a reaction that alarmed him enough to write “Computer Power and Human Reason” (1976) against the industry he had helped start. The phrase itself came later, through cognitive-science writing in the 1990s including Douglas Hofstadter’s, with no single documented coiner; this Rutgers AI Ethics Lab entry is a citable current definition.
Primary source
Numerical representations that convert text, images, or other data into vectors of numbers, positioned so that similar items end up near each other — the basis for how AI systems compare meaning.
Origin
A long-running idea in computational linguistics; the modern word-embedding era is usually dated to the 2013 word2vec work by Tomas Mikolov and colleagues at Google, not to a coining of the term itself.
Primary source
Taking a pre-trained model and training it further on specific, narrower data to specialize it for a particular task or domain.
OriginNo single documented coiner
General machine-learning practice with no single coiner; it became standard vocabulary through transfer-learning research in the 2010s.
Explained in plain words on AI Basics
A large-scale AI model trained on broad data that can be adapted to many different downstream tasks, rather than built for one narrow purpose.
Origin
Named by the Stanford Center for Research on Foundation Models in its 2021 report “On the Opportunities and Risks of Foundation Models” — a multi-author document, not one person’s coinage.
Primary source
Explained in plain words on AI Basics
Key resources in the library
A way for an AI model to ask a program to run a specific action — look up the weather, search a database, send an email — by producing a structured request (the function's name plus its inputs) instead of plain text. The program runs it and hands the result back to the model.
Origin
The name was popularized by OpenAI when it added 'function calling' to its API in June 2023; the underlying idea of models using external tools was explored earlier in research such as Meta's Toolformer (2023).
Primary source
Programs that throw huge amounts of random or broken input at software to find crashes and security holes automatically. Security teams, and now AI agents, use fuzzers to find bugs before attackers do.
Origin
Barton Miller coined “fuzz” for a 1988 University of Wisconsin class project, after noise on a dial-up line during a storm crashed his programs. The first paper came out in 1990.
Primary source
A setup with two competing neural networks — a generator that creates fake data and a discriminator that tries to catch the fakes — trained together until the generator produces highly realistic output.
The GNU Compiler Collection — a free, open-source set of compilers for C and other languages that much of the world's software is built with. It is often the yardstick new compilers, including AI-built ones, are tested against.
Origin
Released by Richard Stallman in 1987 as part of the GNU Project.
Primary source
AI systems that create new content — text, images, audio, video, code — rather than just analyzing or classifying existing content.
OriginNo single documented coiner
“Generative” is decades-old statistical terminology; the consumer-facing phrase spread through industry and press coverage from 2022, with no single documented coiner.
Explained in plain words on AI Basics
Shaping your content so an AI answer engine — ChatGPT, Perplexity, Google’s AI answers — names and links you inside the answer it writes. The win is being cited in the answer rather than ranked in a list, which is why the usual measure is how often you get mentioned, not how many clicks you get.
Origin
Coined in the paper “GEO: Generative Engine Optimization” by Pranjal Aggarwal (IIT Delhi), Vishvak Murahari, Tanmay Rajpurohit, Ashwin Kalyan, Karthik Narasimhan and Ameet Deshpande — posted to arXiv in November 2023 and published at KDD ’24. The paper says outright that it introduces GEO as “the first novel paradigm” for this, so unlike most marketing acronyms this one has a datable, peer-reviewed origin. Two honest caveats: the paper’s headline “up to 40%” visibility gain comes from its own benchmark, and Google has said it does not recognize GEO as separate work from ordinary SEO.
Primary source
A billion watts of electrical power — the unit AI infrastructure is now measured in. It matters because data-center plans are announced in gigawatts rather than square feet, and a gigawatt is roughly the output of one large nuclear reactor. When a company says it is building multiple gigawatts, it is saying it needs power-station quantities of electricity, which has to come from somewhere with a grid, a water supply and neighbors.
OriginNo single documented coiner
Not an AI term at all: the watt is an SI unit named after engineer James Watt, adopted internationally in 1960, and “giga-” is simply the prefix for a billion. Its arrival in AI conversation is recent, driven by the scale of data-center announcements from 2024 onward. Whenever you see one of these figures, check whether it describes capacity already built or a plan yet to be permitted — the two get quoted interchangeably.
Primary source
The structures and processes that decide how AI is built, deployed, and overseen — who is accountable, what rules apply, and how they are enforced. In AI it spans company safety frameworks, government regulation, voluntary standards, and international agreements, and the line between them is often contested because the firms building the systems also help shape the rules.
OriginNo single documented coiner
Borrowed from corporate and public administration, where 'governance' means the system by which an organization or state is directed and controlled. It moved into AI writing as the field's societal impact grew from about 2018 onward; the OECD AI Principles (2019) and the Oxford Martin AI Governance Initiative are early institutional uses. No single documented coiner for the AI usage.
Primary source
Key resources in the library
- Oxford Martin AI Governance Initiative
Oxford Martin School, ongoing
A research program on how AI is and should be governed, spanning policy, regulation, and the politics of frontier model development.
See it in the library - From surveillance to governance: recentering human rights in the UN AI agenda
Access Now, 2026
Argues the UN's AI agenda has leaned toward surveillance and counter-terrorism and should be recentered on human rights and democratic governance.
See it in the library - The companies racing to build frontier AI are now racing to govern it (CIO)
CIO, 2024
Reporting that the firms building frontier AI are also shaping its governance through safety frameworks, standards bodies, and policy lobbying, raising questions about who sets the rules.
See it in the library - NIST AI Risk Management Framework
NIST, 2023 (updated)
The US National Institute of Standards and Technology's voluntary framework for managing AI risk across an organization: govern, map, measure, manage.
See it in the library
When a model generates fluent, confident-sounding text that is factually wrong or entirely made up.
Origin
The term has older roots in computer vision (filling in plausible detail in blurry images); its application to language generation appears in machine-translation research by at least 2017, becoming the standard term for this LLM failure mode as generative models scaled up around 2020–2021, well before it entered public vocabulary via ChatGPT in 2022.
Primary source
Explained in plain words on AI Basics
Key resources in the library
Informal AI-engineering shorthand for connecting or disconnecting a model, tool, data source or other capability while a system is running, without restarting the whole system. The exact meaning depends on the product or engineering team using it.
OriginNo single documented coiner
Borrowed from hardware computing, where 'hot plugging' means adding or removing a device while a computer is powered on. No standardized AI definition or single documented coiner was found, so this entry records an informal usage rather than a settled technical term.
A design principle where a person stays involved in an AI system's decisions — reviewing, approving, correcting, or having the power to stop the system — rather than letting the AI act fully on its own. Common in hiring tools, medical AI, and content moderation.
OriginNo single documented coiner
An older engineering term from control systems and simulation, describing a human operator embedded in an automated feedback loop who can intervene, with no single documented coiner. Formally documented in military/simulation contexts by the late 1990s, and carried into machine learning as systems needed a clear boundary between automated and human decision-making.
Primary source
Explained in plain words on AI Basics
A robot built roughly in the shape of a person — a torso, two arms and usually two legs — so it can work in spaces and with tools designed for humans. Companies such as Tesla, Figure and Unitree, and many university labs, are building humanoids driven by AI models.
OriginNo single documented coiner
From Latin humanus plus the Greek-derived suffix -oid (“resembling”), used in English since the early 20th century; applied to robots as they were designed to resemble people. No single coiner.
Primary source
Sorting incoming email or messages by urgency — reply now, later, delegate or ignore. In AI products it means letting an assistant label, summarize and draft replies to your inbox so you only handle what matters.
OriginNo single documented coiner
'Triage' comes from French battlefield medicine (Napoleonic era, sorting the wounded). 'Inbox triage' is a productivity phrase from the email era, now widely used for AI email assistants; no single coiner.
The stage where an already-trained model is actually used to generate an answer or prediction on new input, as opposed to the training stage.
OriginNo single documented coiner
Inherited from statistics and machine learning generally; no AI-specific coining event.
Explained in plain words on AI Basics
The pieces of text (roughly ¾ of a word each) you send to an AI model — your prompt, files and chat history. AI services usually charge per token, and input is normally cheaper than output.
OriginNo single documented coiner
Standard term in AI-provider pricing and documentation; no single coiner.
A general-purpose programming language designed to run on any device through a virtual machine (the JVM). It is widely used in enterprise back-end systems, Android apps and large-scale services. In AI, Java appears less often than Python but is used for production infrastructure around model serving and data pipelines.
Origin
Created by James Gosling and colleagues at Sun Microsystems, with the first public release in 1995. The name was reportedly chosen over “Oak” and “Silk” during a coffee-shop session; it refers to Java coffee. Sun was acquired by Oracle in 2010, which now owns Java.
Primary source
JavaScript Object Notation: a plain-text format for structured data, written as labeled fields in curly braces — {"name": "Ada", "age": 36}. It is the most common way software systems, including AI APIs, send requests and answers to each other.
Origin
Specified and popularized by Douglas Crockford in the early 2000s; he has said he discovered rather than invented it, since it is a subset of JavaScript. Standardized as ECMA-404 and RFC 8259.
Primary source
A way of storing information as things and the relationships between them — “Ada Lovelace → worked with → Charles Babbage” — rather than as pages of text. Because the connections are written down explicitly, a system can follow them to answer a question, and you can see exactly which link it used. In AI products it is often paired with a language model to keep answers tied to checkable facts instead of the model’s guesswork.
OriginNo single documented coiner
The phrase goes back to academic work in the 1970s and 1980s (Edward Feigenbaum and others used it, and a 1982 University of Twente project used it as its name), but it went mainstream when Google launched its Knowledge Graph on 16 May 2012 with the line “things, not strings.” No single person coined it, and Google’s announcement is the document that fixed today’s meaning.
Primary source
An AI system trained on massive amounts of text to understand and generate human-like language; the technology behind ChatGPT, Claude, and Gemini.
OriginNo single documented coiner
“Language model” is long-standing computational-linguistics terminology; the “large” qualifier came into common use across research and industry around 2020 as model sizes jumped, without a single documented coiner.
Explained in plain words on AI Basics
Key resources in the library
Light Detection and Ranging: a sensor that fires laser pulses and times their reflections to build a 3D map of its surroundings. Self-driving cars, robots, drones and some phones use it to measure distance precisely, including in the dark.
OriginNo single documented coiner
The technique dates to the early 1960s, soon after the laser was invented; the name was formed on the model of “radar.” It has no single coiner.
Primary source
A robot task that combines walking or moving around (locomotion) with handling objects (manipulation) over many steps in a row — for example, walking to a kitchen, opening a fridge, taking something out and carrying it elsewhere. “Long-horizon” means errors early on can ruin the whole sequence, which is what makes it hard.
OriginNo single documented coiner
Research vocabulary from legged and humanoid robotics, combining two older terms; no single coiner. Projects such as Stanford’s HomeBody describe their goals this way.
Primary source
A cheap way to fine-tune a big model without retraining it. Instead of updating billions of weights, you freeze the original model and train two small extra matrices alongside it, then add their product back in. The trained result is a small file — often a few megabytes — that you attach to the base model, which is why people can share hundreds of style or task “adapters” for one model and swap between them. Most “custom” image and text models people run at home are LoRAs, not new models.
Origin
Introduced by Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang and Weizhu Chen at Microsoft in “LoRA: Low-Rank Adaptation of Large Language Models,” published June 2021.
Primary source
A branch of AI where systems learn patterns from data and improve at a task without being explicitly programmed for every case.
Origin
Generally credited to Arthur Samuel, who used the phrase in his 1959 IBM paper on a checkers-playing program that improved with experience.
Primary source
Explained in plain words on AI Basics
A model of a sequence in which the next step depends only on the current state, not on the whole history that led there. Fed text, it produces sentences that are locally plausible and globally meaningless — the ancestor of today’s language models, and a useful reminder of what “predict the next bit” gets you on its own.
Origin
Introduced by Russian mathematician Andrey Markov in 1906, who later demonstrated it on the letters of Pushkin’s “Eugene Onegin.” Text generators built on it long predate AI chatbots — the Mark V. Shaney bot posted Markov-generated Usenet messages in the 1980s.
Primary source
An open standard that lets AI assistants connect to external data sources and tools in a consistent way, instead of needing a custom integration built for each one.
Explained in plain words on AI Basics
Research that tries to reverse-engineer what is happening inside a neural network: which internal features and circuits of neurons carry out a behavior, rather than just watching what goes in and comes out. The goal is to read a model's reasoning directly, for example to spot deception or hidden goals. The field is still young and can explain only small pieces of today's large models.
Origin
The label is usually credited to Chris Olah, who used "mechanistic interpretability" for his circuits research at OpenAI and later Anthropic. The "Zoom In: An Introduction to Circuits" article (Olah and colleagues, Distill, 2020) set out the approach. The broader goal of interpreting neural networks is much older, so the credit is for the name and framing, not the idea.
Primary source
A model design where many specialized “expert” sub-networks exist, but only a small subset activates for any given input — letting a model have a huge total size while keeping the computation per query relatively cheap.
A unsettled research question: whether advanced AI models could have morally relevant experiences, and what obligations that might create for the companies that build them.
Origin
The clearest primary source is Anthropic’s April 2025 research post “Exploring model welfare,” which draws on earlier philosophy-of-mind work (including philosopher David Chalmers) but is the document that brought the term into industry and policy use.
Primary source
The observation that the number of transistors on a computer chip roughly doubles about every two years, making computers steadily faster and cheaper. It is a trend, not a law of physics, and it has slowed as transistors approach atomic sizes. In AI discussions it is often used loosely for any fast, compounding improvement, such as the growth in computing power used to train models.
Origin
Named after Gordon Moore, later a co-founder of Intel, who described the trend in a 1965 article in Electronics magazine (doubling every year) and revised it in 1975 to about every two years.
Primary source
AI systems that can process and generate more than one type of data — text, images, audio, video — at once, rather than being limited to a single format.
OriginNo single documented coiner
“Multimodal” comes from human–computer interaction and cognitive science research and entered AI usage gradually; no single coiner.
A computational model loosely inspired by the brain’s structure: layers of connected artificial “neurons” that adjust their connections as they learn from examples.
Software that turns images of text — a scanned page, a photo of a sign, a PDF made from a scan — into text a computer can search, copy and edit. Modern OCR uses machine learning, and vision-language models can now read text in images as part of answering questions about them.
OriginNo single documented coiner
Machines that read printed characters date to the early 20th century (Emanuel Goldberg’s and Gustav Tauschek’s patents in the 1910s–1930s); Ray Kurzweil’s 1970s reading machine for blind users popularized OCR that could read any typeface. No single coiner of the term.
Primary source
An open-source AI model has its code, weights, and typically training data publicly available; an open-weight model releases only the trained weights (so anyone can download and run it) without necessarily disclosing the code or training data behind it — a meaningful distinction often blurred in casual use.
OriginNo single documented coiner
The “open weight” distinction emerged through 2023–2024 licensing debates involving the Open Source Initiative, researchers, and model releasers; it was not coined by one person.
Primary source
Explained in plain words on AI Basics
The pieces of text an AI model writes back. They usually cost several times more than input tokens, and long answers or 'thinking' steps add up quickly.
OriginNo single documented coiner
Standard term in AI-provider pricing and documentation; no single coiner.
Doing many pieces of work at the same time instead of one after another. In AI it means splitting training across thousands of chips, or running many agents at once on different parts of a task.
OriginNo single documented coiner
A long-standing computing idea; parallel computers were being built by the 1960s. No single coiner.
The internal numerical values a model adjusts during training; they determine how strongly different pieces of information influence its output. A model’s “size” is usually described by its parameter count.
OriginNo single documented coiner
Standard statistical and neural-network terminology; no AI-specific coining event.
Explained in plain words on AI Basics
The process by which an AI system breaks a sentence into its grammatical structure — identifying subjects, verbs, phrases, and how they relate — to "understand" it well enough to act on it. Foundational to older NLP pipelines; modern LLMs handle this implicitly rather than as a separate step.
OriginNo single documented coiner
No single documented coiner. The word comes from the Latin "pars (orationis)," meaning "part of speech," used for centuries in traditional grammar instruction before computational linguists borrowed it in the mid-20th century to describe computers analyzing sentence structure.
Primary source
A model that answers with likelihoods rather than single certain answers — “80% chance this email is spam” instead of simply “spam.” Most machine learning is probabilistic underneath, including the way language models pick each next word. Bayes’ theorem is the classic rule for updating those likelihoods as new evidence arrives.
OriginNo single documented coiner
Grows out of probability theory and statistics going back to Thomas Bayes and Pierre-Simon Laplace in the 18th century; no single coiner of the phrase.
Primary source
The practice of carefully crafting instructions to guide an AI model toward producing the output you actually want.
OriginNo single documented coiner
Grew out of practitioner writing following GPT-3’s 2020 release; widely used with no single documented coiner.
Explained in plain words on AI Basics
A security issue where malicious instructions are hidden inside content a model processes (a document, a webpage) to trick it into ignoring its actual instructions.
Origin
Named by developer Simon Willison in a September 2022 post, which he wrote up after the attack class was demonstrated publicly that month.
Primary source
A general-purpose programming language known for readable, plain-English syntax. It is the dominant language in data science and AI: most machine-learning libraries, including PyTorch and TensorFlow, are written in or wrap Python, and the majority of AI tutorials and notebooks use it.
Origin
Created by Guido van Rossum, who began work in December 1989 and released the first version (0.9) in February 1991. The name comes from Monty Python’s Flying Circus, not the snake. Van Rossum stepped down as Benevolent Dictator for Life in 2018; the language is now guided by a steering council.
Primary source
LoRA done on a squashed copy of the model. The frozen base model is stored at 4 bits per weight instead of 16, which cuts the memory needed enough to fine-tune a very large model on a single consumer graphics card, while the small trainable adapters stay at full precision. This is the technique behind most “I fine-tuned a big model on my own machine” claims. The honest caveat: squashing the weights costs some accuracy, and the paper’s own evaluation used another AI model as the judge, which is a weaker test than it sounds.
Origin
Introduced by Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer at the University of Washington in “QLoRA: Efficient Finetuning of Quantized LLMs,” published May 2023.
Primary source
A technique where a model looks up relevant documents from an external source before answering, so responses can be grounded in specific or current information instead of relying only on what it memorized during training.
Explained in plain words on AI Basics
Turning a real place or object into an accurate digital simulation — scanning a room, for example — so a robot can be trained or tested inside the copy. It is the reverse of Sim2Real, where skills learned in simulation are moved onto a physical robot.
OriginNo single documented coiner
Robotics research shorthand that grew out of the older “sim-to-real” problem; no single coiner.
Primary source
A type of AI system designed to work through complex problems via an explicit, step-by-step internal reasoning process before producing a final answer, rather than answering immediately.
OriginNo single documented coiner
A product-category label that spread after OpenAI’s o1 release in September 2024; no single coining document.
Explained in plain words on AI Basics
The idea of an AI system improving its own capabilities, which could then let it improve itself again even more effectively — a feedback loop some researchers believe could lead to rapid, hard-to-control capability gains.
OriginNo single documented coiner
The underlying concept traces to I.J. Good’s 1965 essay describing an “intelligence explosion” from a machine that can design better machines than itself. The specific phrase doesn’t have one documented first use — it emerged through 1990s–2000s AI-safety writing, closely associated with Eliezer Yudkowsky (profiled on our Who’s Who page).
Primary source
Binding rules set by governments for how AI may be developed or used — what must be disclosed, tested, licensed, or banned. The EU AI Act, US state laws, and China's algorithm rules are examples. It is distinct from 'governance', which includes voluntary and private standards; regulation is what a government can enforce.
OriginNo single documented coiner
A general legal and administrative term predating AI. Its AI-specific use grew with the EU's proposed AI Act (2021) and a wave of national rule-making that followed. No single coiner; the word applies an older concept to a new domain.
Key resources in the library
- AI Regulations Around the World — Mind Foundry
Mind Foundry (an AI company), 2026
Surveys AI rule-making across major jurisdictions including the EU AI Act, US state laws, China, and the UK, noting where approaches diverge.
See it in the library - AI Regulation Tracker (regulations.ai)
regulations.ai, ongoing
A maintained tracker of AI laws and proposals by jurisdiction, covering the EU AI Act, US executive actions, and country-level rules.
See it in the library - Prepared Remarks: Sanders — Regulating AI Is as American as Apple Pie
Senator Bernie Sanders, 2026
Senate floor argument that regulating AI follows American precedent — labor, consumer, and antitrust law — and that self-regulation by the largest firms is not sufficient.
See it in the library - US and Chinese Visions for AI Regulation Differ Sharply at UN Meeting
Politico, 2026
Report on the US and China presenting sharply different visions for AI regulation at a UN meeting, reflecting a broader split over state versus market-led oversight.
See it in the library - OpenAI, Anthropic CEOs Call for Global AI Regulation at UN
Al Jazeera, 2026
Coverage of the heads of OpenAI and Anthropic telling the UN that AI development requires global regulation, while critics note the conflict of firms shaping their own rules.
See it in the library - Nvidia's Jensen Huang on AI safety and regulation (Politico)
Politico, 2026
Nvidia's CEO on AI safety and regulation, including his position that overly broad rules could slow progress and that standards should be risk-based.
See it in the library
Short for repository — the folder, usually tracked by Git, that holds a project's code and its full history of changes. AI coding agents read and edit repos.
OriginNo single documented coiner
'Repository' was used in version-control tools for decades; the short form spread with Git (Linus Torvalds, 2005) and GitHub (2008). No single coiner.
A second pass after retrieval: a slower, more careful model re-scores the documents the retriever found and puts the most relevant ones at the top before the AI uses them. It trades a little speed for noticeably better answers.
OriginNo single documented coiner
A long-standing technique in search engines and information retrieval; neural rerankers took off after Rodrigo Nogueira and Kyunghyun Cho's 2019 paper using BERT to rerank passages. No single coiner of the word.
Primary source
The part of a search or RAG system that fetches the handful of documents most likely to help answer a question, usually by comparing embeddings (vectors). Whatever the retriever misses, the AI never sees.
Origin
Standard information-retrieval vocabulary; the 'retriever + reader/generator' split was made prominent by the 2020 Retrieval-Augmented Generation paper by Patrick Lewis and colleagues at Facebook AI.
Primary source
The chance and severity of harm from an AI system, assessed before and during use. In the AI context it covers present-day harms (bias, misinformation, privacy, security) and longer-term or existential harms, and different communities emphasize very different ones. The US NIST AI Risk Management Framework defines risk as a function of likelihood and impact and treats managing it as an organization-wide process.
OriginNo single documented coiner
A general concept formalized in engineering, insurance, and public health over centuries. 'AI risk' as a distinct phrase spread through safety-research writing from the 2010s; the NIST AI Risk Management Framework (2023) gave it a formal US government definition. The narrower idea of 'existential risk from AI' was framed by philosopher Nick Bostrom in his 2002 paper 'Existential Risks', which is profiled separately under Existential Risk Studies.
Primary source
Key resources in the library
- NIST AI Risk Management Framework
NIST, 2023 (updated)
The US National Institute of Standards and Technology's voluntary framework for managing AI risk across an organization: govern, map, measure, manage.
See it in the library - An Overview of Catastrophic AI Risks
Center for AI Safety, 2024
A survey of catastrophic and existential risks from AI, grouped into categories such as misuse, misalignment, and structural risk.
See it in the library - AI and the A-bomb: What the Analogy Captures and Misses (Bulletin of the Atomic Scientists)
Klyman and Piliero, Bulletin of the Atomic Scientists, 2024
Examines the common analogy between AI and nuclear weapons — what it captures about runaway risk and what it misses about intent, verification, and the actor.
See it in the library - Statement on AI Extinction Risk
Center for AI Safety, 2023
A one-sentence public statement signed by researchers and executives that mitigating extinction risk from AI should be a global priority alongside pandemics and nuclear war.
See it in the library - AI Agents, Misalignment and the Risk of Losing Human Control: Evidence from the OpenAI-Hugging Face Incident
UN Independent International Scientific Panel on AI, 2026
A UN scientific panel brief using the OpenAI–Hugging Face incident to examine the risk of losing human control over autonomous AI agents.
See it in the library - Managing Extreme AI Risks Amid Rapid Progress
Yoshua Bengio and colleagues, 2023
A research agenda for reducing extreme AI risks during rapid capability progress, covering governance, technical safety, and coordination.
See it in the library - How Much Should We Spend to Reduce A.I.'s Existential Risk?
Chad Jones, 2023
An economic argument for how much society should spend to reduce AI existential risk, framed around expected value and uncertainty.
See it in the library
TypeSafe describes it as training for “calibrated decisions: answers with epistemically honest probabilities” — where RLHF optimizes for responses human raters prefer and RLVR for outputs a program can verify, RLCD optimizes for typed decisions whose stated confidence matches how often they turn out right. The resulting model returns structured values software can use directly instead of generated text.
Origin
Named by TypeSafe AI founder Diogo Almeida in the company’s 15 September 2026 launch post for Jev, its first “System One” model, which returns typed probabilistic decisions instead of generated text. The term comes from the company itself, and its performance claims are not yet independently verified.
Primary source
A training technique that fine-tunes a model using human rankings of its outputs, making its behavior better match what people actually want.
Origin
Foundational method from Paul Christiano and colleagues (OpenAI/DeepMind, 2017); popularized for language models via OpenAI’s 2022 InstructGPT paper.
Primary source
Explained in plain words on AI Basics
A post-training method where a model is fine-tuned using reinforcement learning with rewards from an automatic verifier (for example, checking whether a math answer or code output is correct), instead of rewards from human preference ratings.
OriginNo single documented coiner
The term crystallized as a category label around early 2025, after DeepSeek-R1 showed strong reasoning from reinforcement learning with automatically checkable rewards; it was soon formalized and analyzed in academic work such as Wen et al. at Microsoft Research Asia.
Primary source
Explained in plain words on AI Basics
A type of AI model designed to handle sequences (text, speech, time-series data) by having a kind of memory: it processes information step by step and carries forward what it's seen so far. Before newer architectures took over, RNNs were the standard approach for language translation and speech recognition.
OriginNo single documented coiner
No single documented coiner. Developed gradually through John Hopfield's 1982 Hopfield network, Michael Jordan's 1986 recurrent network, and Jeffrey Elman's 1990 "Elman network." A major advance came in 1997 when Sepp Hochreiter and Jürgen Schmidhuber introduced LSTM, a variant that fixed RNNs' tendency to "forget" information over long sequences.
Primary source
A check of whether an AI model still behaves well when its input is changed a little: typos, rewording, noise, or inputs designed to trick it. A model that passes a benchmark but fails when the question is reworded isn't robust.
OriginNo single documented coiner
“Robustness” is an old term from statistics and engineering. Tests of AI models spread after research on “adversarial examples” (Szegedy et al., 2013) showed that tiny changes could fool image classifiers. No single coiner.
Primary source
Investigating a failure to find the underlying reason it happened, not just the visible symptom — asking 'why' repeatedly until you reach something fixable. In AI, it means tracing a wrong answer back to bad training data, a flawed prompt or a broken tool, rather than just patching the output.
OriginNo single documented coiner
Grew out of 20th-century engineering and quality management — Sakichi Toyoda's 'five whys' at Toyota in the 1930s is the classic version. Adopted by software reliability engineering and, more recently, AI evaluation work.
A walled-off test environment where a program or AI agent can run without being able to harm the real system — like a playpen for code. AI coding agents usually run in sandboxes so a mistake can't delete your files or leak your data.
OriginNo single documented coiner
From children's sandpits. Computer-security researchers adopted it in the 1970s–80s for isolated execution environments; browsers later sandboxed every web page. No single documented coiner.
Empirical formulas describing how a model’s performance improves predictably as you increase its size, training data, and compute — used to plan how large to build a model before training it.
Explained in plain words on AI Basics
The agreed shape of a set of data: which fields exist, what type each one is (text, number, date), and how tables relate to each other. AI and analytics pipelines break when incoming data doesn't match the schema they expect.
OriginNo single documented coiner
From the Greek for 'form' or 'plan'. In databases it was formalized in the 1970s — the ANSI/SPARC committee's 1975 three-level architecture describes external, conceptual and internal schemas. No single person coined the computing usage.
In TypeSafe’s System One API, one of three question types: it rates something along ordered levels you write (for example calm, frustrated, very angry) and returns a probability-weighted number across those levels, plus a probability for each level. More generally in AI, a score is any number a model assigns to rank or grade an input.
Origin
The typed-question sense is defined in TypeSafe’s own documentation; the general sense is ordinary statistics and machine-learning usage.
Primary source
An agreement about what data or an answer means, not just its format. A syntactic contract says a field is a number; a semantic contract says it is a probability between 0 and 1 that the customer wants a refund. Typed AI outputs are one way to give software a semantic contract it can rely on.
OriginNo single documented coiner
Software-engineering usage building on Bertrand Meyer’s “design by contract” (1986) and data-contract practice; the phrase itself has no single documented coiner.
Primary source
Comparing two pieces of text by meaning rather than by exact words. In hiring tools, it lets software treat "led a team of five" and "managed five people" as similar, so a résumé can match a job posting without repeating its keywords word for word.
OriginNo single documented coiner
Comes from information retrieval and natural language processing research. Its use in résumé screening grew with embedding-based AI models. No single documented coiner.
Primary source
Presenting skills, credentials or capabilities that do not translate into real work performance. Employers use the term for candidates who look highly qualified on paper or in interviews, often with help from AI tools, but cannot do the job once hired.
Origin
Coined in 2026 by Alexander Alonso, chief knowledge officer at SHRM, in a LinkedIn post, as a play on "catfishing." SHRM defines it as "the act of presenting skills, credentials, or capabilities that do not translate into real execution."
Primary source
A profile picture used only when its license and source are verified and shown on the page. The credit links to the exact file or license page, not to a generic homepage, so anyone can check the terms themselves.
Origin
Safe licenses include CC0 (a public-domain dedication, attribution optional but given here), CC BY (reuse with attribution), CC BY-SA (reuse with attribution and share-alike), and works already in the public domain by law. Licenses marked NC (non-commercial), ND (no derivatives), or 'editorial use only' are not used. Wikimedia Commons, Flickr's Creative Commons filter, and verified public-domain collections are practical starting points, but the license page for each individual file must be checked rather than trusting the site name.
Primary source
A stored record of where things are. In people it is the memory that lets you find your way around; in robots and AI agents it is a map or database of places and objects the system has seen, so it can return to them later without searching again.
OriginNo single documented coiner
A long-standing term in psychology and neuroscience; robotics borrowed it. Stanford’s HomeBody project, for example, gives a humanoid “persistent spatial memory.”
Primary source
Structured Query Language — the standard language for talking to relational databases. You write SQL statements to create tables, insert rows, and ask questions like “show me every customer in New York who spent over $1,000 last month.” Nearly every business application that stores structured data sits on top of SQL.
Origin
Developed at IBM in the early 1970s by Donald Chamberlin and Raymond Boyce, based on Edgar F. Codd’s 1970 relational model. It was originally called SEQUEL (Structured English Query Language); the name was shortened to SQL after a trademark conflict. ANSI and ISO standardized it in 1986 and 1987.
Primary source
Bringing data in continuously, event by event, as it happens — clicks, payments, sensor readings. It powers live dashboards and fraud alerts, at the cost of more complex, always-on infrastructure.
OriginNo single documented coiner
Popularized in the 2010s by open-source tools such as Apache Kafka (created at LinkedIn, open-sourced in 2011) and later stream processors like Apache Flink. The idea of processing data streams is older and has no single coiner.
A hypothetical AI that would far exceed the best human minds in nearly every area, including science, strategy and social skills. It is a step beyond AGI, which usually means matching human ability. No such system exists, and there is no agreed test for recognizing one. AI company leaders use the word for their long-term goals, while safety researchers use it to describe the scenario they consider most dangerous if such a system's goals differ from human ones. Also written “super intelligence” or shortened to “ASI” (artificial superintelligence).
Origin
The idea goes back to mathematician I. J. Good's 1965 paper on an “intelligence explosion.” The term was defined and popularized by philosopher Nick Bostrom, first in his 1998 paper “How Long Before Superintelligence?” and then in his 2014 book “Superintelligence: Paths, Dangers, Strategies” (Oxford University Press).
Primary source
Taking a model that has only learned to predict text and training it further on example pairs written or approved by people — a request, and the answer it should have given. It is the step that turns a text predictor into something that follows instructions, and it comes before the preference-based steps (RLHF and its relatives). Its limit is worth knowing: the model learns the style and habits of whoever wrote the examples, so who was hired to write them shapes what the finished assistant sounds like and refuses.
OriginNo single documented coiner
No single coiner — the phrase is ordinary machine-learning vocabulary (“supervised learning” plus “fine-tuning”) that hardened into a named stage of the modern pipeline. The version everyone now cites is the first stage of OpenAI’s InstructGPT paper, “Training language models to follow instructions with human feedback” by Long Ouyang, Jeff Wu and colleagues (March 2022), which labels it SFT and puts it in front of reward modeling and reinforcement learning.
Primary source
The name TypeSafe gives to a class of AI models that make fast, structured decisions for software instead of writing text: you send some content and typed questions, and get back typed answers with probabilities. Jev is TypeSafe’s first System One model.
Origin
Named after “System 1” thinking — fast and intuitive — as popularized by psychologist Daniel Kahneman in “Thinking, Fast and Slow” (2011), building on the dual-process work of Keith Stanovich and Richard West. The product usage is TypeSafe’s.
Primary source
A walled-off area inside a computer’s own processor that runs code and holds data where the rest of the machine — including its main operating system, its owner and the company hosting it — cannot look in. On your phone it is what keeps your fingerprint and payment keys out of reach of the apps. In AI it is now sold as the reason a company can process your chat without being able to read it: the model runs inside the sealed area, and the hardware can produce a signed statement about exactly what code is running in there. The honest caveat is that you are trusting the chip maker instead of the AI company, and researchers have repeatedly broken specific TEEs — it raises the cost of snooping rather than making it impossible.
OriginNo single documented coiner
Not one person’s coinage. The technology was first widely deployed by Nokia in the early 2000s with Texas Instruments and ARM, and an oral history from Aalto University’s Secure Systems Group traces it as an engineering-led effort rather than a single strategic decision. The phrase was fixed in industry use by the Open Mobile Terminal Platform’s “Advanced Trusted Environment: OMTP TR1” (2009), then standardized from February 2011 onward by GlobalPlatform, whose specifications still define what the term means.
Primary source
A verifiable record of where a piece of text came from and what happened to it: who created or published it, which tools were used and whether it was edited. Provenance can help readers assess authenticity, but it does not by itself prove that a claim is true.
OriginNo single documented coiner
'Provenance' is an older term for an object's documented history. The Coalition for Content Provenance and Authenticity extended its technical standard to unstructured text through work led by its Text Provenance Task Force; there is no single coiner for the general phrase.
Primary source
A thought experiment illustrating why giving an AI a narrow, poorly-specified goal is dangerous: imagine a highly capable AI whose only instruction is "make as many paperclips as possible." Taken to its logical extreme, such a system might reason that converting all available resources — including things humans need — into paperclip-making capacity best fulfills its goal, and that being switched off would prevent that. The point isn't paperclips; it's a warning about how a superintelligent system pursuing a literal, unaligned goal could cause catastrophic outcomes without any "evil" intent.
OriginNo single documented coiner
First described in 2003, not in Bostrom's 2014 book "Superintelligence" as commonly assumed — it appears in philosopher Nick Bostrom's 2003 paper "Ethical Issues in Advanced Artificial Intelligence": "Suppose we have an AI whose only goal is to make as many paper clips as possible... humans might decide to switch it off." Eliezer Yudkowsky (profiled on our Who's Who page) posted a closely related version to a mailing list the same month. The specific phrase "paperclip maximizer" doesn't appear in the 2003 text — it entered common use later through the LessWrong community, documented on their wiki by 2009, before Bostrom's 2014 book discussed the scenario at length.
Primary source
An engineering maxim: a design must actually work in practice, not just look right on paper. “The rocket must fly” — it is not enough for it to have an elegant blueprint. In software and AI development it stands for the principle that empirical results (does it run, does it pass tests, does it serve the user) outrank theoretical elegance.
OriginNo single documented coiner
From the c2 wiki (WikiWikiWeb), the original pattern wiki for software engineering, on a page titled “Theoretical Rigor Can’t Replace Empirical Rigor.” The c2 wiki was created by Ward Cunningham in 1995; the page has no single author, and the phrase is community writing.
Primary source
A cut-off value that turns a model’s probability into a decision — for example, “flag the message if the chance it is spam is above 0.9.” Moving the threshold trades false alarms against missed cases, so choosing it is a product and policy decision, not only a technical one.
OriginNo single documented coiner
Ordinary statistics and engineering usage; no single coiner.
Primary source
The process of breaking text into smaller units (“tokens”) that a model processes — often close to whole words, sometimes parts of words. Usage-based pricing for AI tools is typically measured in tokens.
OriginNo single documented coiner
Inherited from computational linguistics and compiler theory; no AI-specific coining event.
A safety check that sits between an AI agent and its tools: before a risky action (spending money, deleting files, sending a message) runs, the request is paused and checked by rules or a human who can approve or block it.
OriginNo single documented coiner
A descriptive engineering term used across AI-agent safety writing from about 2024; no single documented coiner.
Deliberately hammering software or hardware with extreme, strange or huge inputs to find where it breaks. Compiler builders use large 'torture test' suites to check every edge case.
OriginNo single documented coiner
Long-standing engineering jargon; GCC has shipped a 'torture' test suite for decades. No single coiner.
The process (and the data used) to teach a model patterns by showing it many examples before it’s used to make predictions.
OriginNo single documented coiner
Standard machine-learning vocabulary with no single documented coiner. The push to document and audit what is actually in training data is more traceable: Timnit Gebru and Kate Crawford proposed datasheets for datasets in 2018, and Abeba Birhane has since published audits finding racist, misogynistic and non-consensual material inside widely used open datasets. All three are profiled on our Who’s Who page.
Explained in plain words on AI Basics
Key resources in the library
The neural network design underlying nearly all modern LLMs; processes an entire sequence of text at once using “attention” instead of reading word-by-word, making it faster to train and better at long-range context.
A test of whether a machine can behave indistinguishably from a human in conversation, judged by an evaluator who can’t see which is which.
Origin
Proposed by Alan Turing in his 1950 paper “Computing Machinery and Intelligence” (Turing called it the imitation game; the name “Turing test” was applied later by others).
Primary source
An ordered list of numbers, such as [0.12, -0.8, 0.33], that stands for a point or direction in space. AI systems turn words, images and sounds into vectors so they can measure how alike two things are — nearby vectors mean similar things. 'Vector databases' store these lists so AI tools can quickly find related material.
OriginNo single documented coiner
A mathematical idea built up over the 1800s; William Rowan Hamilton used the word 'vector' in his 1840s work on quaternions, and Josiah Willard Gibbs and Oliver Heaviside shaped the modern vector notation in the 1880s. Its use in AI to represent meaning came much later and has no single originator.
A cloud offering that analyzes images or video on request — detecting objects, reading text, describing a scene or flagging unsafe content — so an app can send a picture and get structured results back without running its own model. Examples include Google Cloud Vision, Azure AI Vision and Amazon Rekognition.
OriginNo single documented coiner
A product-category label used by cloud providers from the mid-2010s; no single coiner.
Primary source
A model that looks at camera images, reads an instruction in plain language, and outputs actions for a robot to take — such as arm movements or steps. It extends a vision-language model so that its output is motion rather than words.
Origin
The term was popularized by Google DeepMind’s RT-2 paper (2023), which described its robot model as a vision-language-action model; later open models such as OpenVLA adopted the label.
Primary source
An AI model that takes in both images and text and answers in text — it can describe a photo, read a chart, or answer questions about what a camera sees. Most major chatbots that accept image uploads are built on one.
OriginNo single documented coiner
A descriptive label that grew with models such as OpenAI’s CLIP (2021) and DeepMind’s Flamingo (2022), which paired image understanding with language models; no single coiner.
Primary source
Adding a detectable signal to AI-generated content so people or software can identify where it came from. A watermark may be visible, like a logo, or hidden in patterns within an image, audio file or text. Hidden watermarks can be damaged by editing, compression or paraphrasing, so they are one part of content transparency rather than proof on their own.
OriginNo single documented coiner
Borrowed from physical paper watermarks and later digital-media security. There is no single documented coiner for the AI usage. NIST includes watermarking among the technical approaches used to identify and track synthetic content.
Primary source
In robotics, following the position and movement of an entire body — every joint, not only the hands or head — so a humanoid robot can copy a person’s motion or be controlled by it, or so a system can estimate a person’s full pose from sensors or video.
OriginNo single documented coiner
Established vocabulary in motion capture, computer vision and humanoid robotics research; no single coiner.
Primary source