AI glossary
AI concepts explained simply: machine learning, deep learning, LLMs, NLP, RAG, embeddings, fine-tuning, MCP…
The basics
- Artificial intelligence (AI)
- The set of techniques that let a machine perform so-called intelligent tasks: understanding text, recognizing an image, making a decision.
- Machine learning
- A branch of AI where the machine learns from examples instead of being programmed rule by rule. For instance, from thousands of emails labelled "spam" or "normal", it learns to sort them on its own. Classic methods (regression, decision trees) are mostly used on tabular data.
- Deep learning
- A branch of machine learning that uses neural networks with many layers. It needs a lot of data and compute, but excels at text, images and sound. All recent models (LLMs, image generators, speech recognition) rely on deep learning.
- Neural network
- A mathematical model inspired by the brain, made of connected layers of "neurons". By adjusting the weights of these connections during training, the network learns to turn an input (text, image) into an output (an answer, a category).
- Model
- The result of training: a file of parameters (the weights) that can perform a task. You download it, call it through an API, or run it yourself. See the guide →
The main fields
- NLP (natural language processing)
- Everything about text: writing, summarizing, translating, classifying, extracting information, answering questions.
- Computer vision
- Everything about images and video: recognizing objects, classifying a photo, detecting faces, generating images.
- Multimodal
- A model that handles several kinds of data at once: describing an image in text, answering a question about a photo, generating a video from text.
- Generative AI
- Models that produce new content (text, image, video, sound, code) rather than only classifying or predicting.
Language models
- LLM (large language model)
- A deep learning model trained on huge amounts of text to predict how a text continues. It powers Claude, GPT, Mistral or Llama. See the guide →
- Token
- The piece of text an LLM reads, often part of a word. Model limits (context size) and prices are counted in tokens.
- Context window
- How much text (in tokens) a model can read at once: the question, instructions, documents and answer.
- Hallucination
- A wrong answer given confidently by a model that makes things up instead of saying it doesn't know. RAG and guardrails reduce it.
- Inference
- Using an already trained model to get an answer, as opposed to training. See the guide →
- Quantization
- Lowering the precision of a model's numbers (e.g. from 16 to 4 bits) so it uses less memory and runs on smaller hardware, at a slight cost in quality. See the guide →
Using a model
- Prompt
- The text sent to the model to ask for something. A reusable prompt contains variables, such as {{product}}. See the guide →
- Context (system instruction)
- Instructions and examples given to the model before the question, to set its role, tone and rules, without retraining it. See the guide →
- Few-shot
- Giving a few examples (input and expected output) in the context to show the model what you want. See the guide →
- RAG (retrieval-augmented generation)
- Having a model answer from your documents: relevant passages are retrieved, then given to it with the question. See the guide →
- Embeddings
- Lists of numbers (vectors) that represent the meaning of a text. Texts with similar meaning get close vectors: the basis of semantic search and RAG. See the guide →
- Structured output
- Forcing the model to answer in a precise format, usually JSON described by a schema. See the guide →
Acting and tooling
- Agent
- An AI assistant that chains several steps and uses tools to reach a goal, instead of answering in one go. See the guide →
- MCP (Model Context Protocol)
- An open standard to plug tools and data (files, databases, software) into an AI assistant, compatible with Claude, Cursor and others. See the guide →
- Workflow
- An automated chain of steps where AI takes part, built in a tool such as n8n or LangGraph. See the guide →
Training and adapting
- Dataset
- A set of examples used to train, adapt or evaluate a model. See the guide →
- Fine-tuning
- Retraining an existing model on your own examples to specialize it in a domain. See the guide →
- LoRA
- A low-cost fine-tuning method that trains only a small part of the model. QLoRA applies it to a quantized model. See the guide →
- Evaluation (benchmark)
- A set of tests to measure the quality of a model, a prompt or an agent. See the guide →
Reliability and compliance
- Guardrail
- A rule that controls what goes into or out of an AI assistant: blocking personal data, refusing off-topic requests. See the guide →
- Prompt injection
- An attack that hides instructions in a message or document to turn the model against its rules. See the guide →
- Red teaming
- Deliberately attacking an AI system to find its weaknesses before users do. See the guide →
- AI Act
- The European regulation on artificial intelligence: it sets transparency, documentation and oversight obligations depending on the system's risk level. See the guide →
- GDPR
- The European regulation on personal data protection, which also applies to data sent to an AI or used to train it.