Corpus: AI resource definitions (English)
Definitions of every AI resource type on Contexte Tech, one chunk each with the guide URL: a small English corpus ready for RAG.
Découpage : 1 resource type = 1 chunk
Aperçu
| id | Texte | Source |
|---|---|---|
context | A context is a set of instructions and examples given to the AI model before the user's question. It steers the answer without retraining anything: this is in-context learning. | https://contextetech.com/en/docs/contexts |
prompt | A prompt is a reusable text template with variables in double braces, for example {{product}}. Fill in the variables, then send it to the model. | https://contextetech.com/en/docs/prompts |
dataset | A dataset is a set of examples in JSONL format (one JSON line per example) used to train or evaluate a model. | https://contextetech.com/en/docs/datasets |
lora | LoRA adapts an existing model to a task without fully retraining it: only small adapters are trained. A LoRA page gathers all the settings of that training. | https://contextetech.com/en/docs/lora |
tool | An MCP (Model Context Protocol) server gives an AI assistant new abilities: read files, query a database, call a service. The same server works with Claude, Cursor and other assistants. | https://contextetech.com/en/docs/mcp-servers |
model | A model is an AI model published by a verified Lab: fine-tune, merge or quantized version. The page describes the model and links to its download. | https://contextetech.com/en/docs/models |
agent | An agent is an AI assistant that chains steps to complete a task: it follows instructions, uses MCP servers and contexts, and respects guardrails. | https://contextetech.com/en/docs/agents |
skill | A skill is an instruction pack (SKILL.md file) that an assistant like Claude loads only when the task calls for it. It teaches a precise way of doing things. | https://contextetech.com/en/docs/skills |
eval | An evaluation is a set of test cases (input and expected answer) with a scoring method. It measures whether a model, prompt, context or agent does its job. | https://contextetech.com/en/docs/evals |
rag | A RAG (retrieval-augmented generation) pipeline lets a model answer from your documents: they are chunked and embedded, then for each question the useful excerpts are retrieved and given to the model. | https://contextetech.com/en/docs/rag |
harness | A harness is the environment that runs an agent: its instructions file (e.g. CLAUDE.md), the tools it may use, what is denied, hooks and connected MCP servers. | https://contextetech.com/en/docs/harness |
workflow | A workflow is a chain of steps where AI takes part: fetch data, have a model analyse it, then send the result elsewhere. It is built in a tool such as n8n, LangGraph, Make or Dify, then exported as JSON. | https://contextetech.com/en/docs/workflows |
guardrail | A guardrail is a set of rules that controls what goes into or out of an AI assistant: block personal data, detect prompt injection, refuse off-topic requests. | https://contextetech.com/en/docs/guardrails |
embedding | An embedding model turns a text into a list of numbers (a vector). Texts with similar meaning give close vectors: it is the basis of semantic search and RAG. | https://contextetech.com/en/docs/embeddings |
inference | An inference config is the set of settings needed to run a model on your own machine: the runtime (Ollama, vLLM, llama.cpp), quantization, context size and launch command. | https://contextetech.com/en/docs/inference |
stack | A stack is a set of AI services that work together and start with one command: for example a model, a vector database and a chat interface, described in a docker-compose file. | https://contextetech.com/en/docs/stacks |
redteam | Red teaming means deliberately attacking an AI assistant to find its weaknesses before users do: rule bypass, data leaks, prompt injection, promises it should not make. | https://contextetech.com/en/docs/red-teaming |
aiact | An AI Act file is a documentation template to fill in to comply with the EU AI Act: transparency notice, training data description, risk assessment or system register. | https://contextetech.com/en/docs/ai-act-files |
schema | An output schema is a JSON Schema describing exactly the shape of the expected answer: which fields, which types, which are required. Recent models can follow it (structured outputs). | https://contextetech.com/en/docs/schemas |
corpus | A corpus is a set of documents already split into chunks, each with its text and source. It is ready to embed for RAG, and every answer can cite where it comes from. | https://contextetech.com/en/docs/corpus |
Coût estimé
Tokens d'entrée par appel : ≈ 3 028
| Modèle | Par appel | Pour 1 000 appels |
|---|---|---|
| Mistral Large · Mistral AI | 0,00133 € | 1,33 € |
| DeepSeek Flash · DeepSeek | 0,00080 € | 0,80 € |
| DeepSeek V4 Pro · DeepSeek | 0,00352 € | 3,52 € |
| GPT-6 Luna · OpenAI | 0,00027 € | 0,27 € |
| GPT-6 Sol · OpenAI | 0,00533 € | 5,33 € |
| Claude Sonnet 5.5 · Anthropic | 0,00533 € | 5,33 € |
| Claude Opus 5.5 · Anthropic | 0,01 € | 10,67 € |
| Modèle open source hébergé chez soi (Ollama, vLLM) | 0 € d'API (coût du serveur seulement) | |
Coût d'entrée seulement, sans la réponse du modèle. Prix publics des fournisseurs (Mistral AI, DeepSeek, OpenAI, Anthropic), tarif standard en dollars converti en euros au taux BCE du 29 septembre 2026. Estimation : 1 token ≈ 3,6 caractères.
Commun aux 3 régions
Contenu créé par l'auteur, sans source extérieure.
Usage commercial autorisé · citer l'auteur · modifications permises.
Texte ou données seuls : aucun accès aux fichiers, au réseau ni aux commandes.
Hébergé en France (Scaleway, Paris). Aucun transfert hors UE.
Union européenne · 5/5
Aucune donnée personnelle.
Pas de données d'entraînement : pas de résumé à fournir.
États-Unis · 5/5
Aucune donnée personnelle.
Pas de données d'entraînement : pas de documentation à publier.
Chine · 5/5
Aucune donnée personnelle.
Pas de données d'entraînement ; les contenus générés par IA diffusés en Chine doivent être marqués (2025).
Synthèse indicative au 30/09/2026, sans valeur de certification juridique. Méthode et sources →
Statistiques
J'aime : 0 · Commentaires : 0
Cliquez sur un compteur pour l’afficher ou le masquer, survolez le graphique pour le détail. Une vue par visiteur et par jour, robots exclus ; décompte commencé le 30 septembre 2026.
- Identifiant
- contextetech--ai-resource-definitions-en
- Type
- Corpus
- Fichier
- ai-resource-definitions-en.corpus.jsonl
- Format
- JSON Lines (application/jsonl)
- Taille
- 4 852 octets (4,7 Ko)
- Encodage
- UTF-8
- Tokens estimés
- ≈ 1 348
- Langue
- Anglais
- Licence
- MIT
- Usage commercial
- Autorisé
- Données personnelles
- Aucune
- Version
- 1.0
- Publié le
- 1 octobre 2026 à 18:55
- Mis à jour le
- 1 octobre 2026 à 18:55
- Stockage
- en base
- Utilisations
- 0
- J'aime
- 0
ai-resource-definitions-en.corpus.jsonl
{"id":"context","text":"A context is a set of instructions and examples given to the AI model before the user's question. It steers the answer without retraining anything: this is in-context learning.","source":"https://contextetech.com/en/docs/contexts"}
{"id":"prompt","text":"A prompt is a reusable text template with variables in double braces, for example {{product}}. Fill in the variables, then send it to the model.","source":"https://contextetech.com/en/docs/prompts"}
{"id":"dataset","text":"A dataset is a set of examples in JSONL format (one JSON line per example) used to train or evaluate a model.","source":"https://contextetech.com/en/docs/datasets"}
{"id":"lora","text":"LoRA adapts an existing model to a task without fully retraining it: only small adapters are trained. A LoRA page gathers all the settings of that training.","source":"https://contextetech.com/en/docs/lora"}
{"id":"tool","text":"An MCP (Model Context Protocol) server gives an AI assistant new abilities: read files, query a database, call a service. The same server works with Claude, Cursor and other assistants.","source":"https://contextetech.com/en/docs/mcp-servers"}
{"id":"model","text":"A model is an AI model published by a verified Lab: fine-tune, merge or quantized version. The page describes the model and links to its download.","source":"https://contextetech.com/en/docs/models"}
{"id":"agent","text":"An agent is an AI assistant that chains steps to complete a task: it follows instructions, uses MCP servers and contexts, and respects guardrails.","source":"https://contextetech.com/en/docs/agents"}
{"id":"skill","text":"A skill is an instruction pack (SKILL.md file) that an assistant like Claude loads only when the task calls for it. It teaches a precise way of doing things.","source":"https://contextetech.com/en/docs/skills"}
{"id":"eval","text":"An evaluation is a set of test cases (input and expected answer) with a scoring method. It measures whether a model, prompt, context or agent does its job.","source":"https://contextetech.com/en/docs/evals"}
{"id":"rag","text":"A RAG (retrieval-augmented generation) pipeline lets a model answer from your documents: they are chunked and embedded, then for each question the useful excerpts are retrieved and given to the model.","source":"https://contextetech.com/en/docs/rag"}
{"id":"harness","text":"A harness is the environment that runs an agent: its instructions file (e.g. CLAUDE.md), the tools it may use, what is denied, hooks and connected MCP servers.","source":"https://contextetech.com/en/docs/harness"}
{"id":"workflow","text":"A workflow is a chain of steps where AI takes part: fetch data, have a model analyse it, then send the result elsewhere. It is built in a tool such as n8n, LangGraph, Make or Dify, then exported as JSON.","source":"https://contextetech.com/en/docs/workflows"}
{"id":"guardrail","text":"A guardrail is a set of rules that controls what goes into or out of an AI assistant: block personal data, detect prompt injection, refuse off-topic requests.","source":"https://contextetech.com/en/docs/guardrails"}
{"id":"embedding","text":"An embedding model turns a text into a list of numbers (a vector). Texts with similar meaning give close vectors: it is the basis of semantic search and RAG.","source":"https://contextetech.com/en/docs/embeddings"}
{"id":"inference","text":"An inference config is the set of settings needed to run a model on your own machine: the runtime (Ollama, vLLM, llama.cpp), quantization, context size and launch command.","source":"https://contextetech.com/en/docs/inference"}
{"id":"stack","text":"A stack is a set of AI services that work together and start with one command: for example a model, a vector database and a chat interface, described in a docker-compose file.","source":"https://contextetech.com/en/docs/stacks"}
{"id":"redteam","text":"Red teaming means deliberately attacking an AI assistant to find its weaknesses before users do: rule bypass, data leaks, prompt injection, promises it should not make.","source":"https://contextetech.com/en/docs/red-teaming"}
{"id":"aiact","text":"An AI Act file is a documentation template to fill in to comply with the EU AI Act: transparency notice, training data description, risk assessment or system register.","source":"https://contextetech.com/en/docs/ai-act-files"}
{"id":"schema","text":"An output schema is a JSON Schema describing exactly the shape of the expected answer: which fields, which types, which are required. Recent models can follow it (structured outputs).","source":"https://contextetech.com/en/docs/schemas"}
{"id":"corpus","text":"A corpus is a set of documents already split into chunks, each with its text and source. It is ready to embed for RAG, and every answer can cite where it comes from.","source":"https://contextetech.com/en/docs/corpus"}
Charger le corpus (Python)
import json
chunks = [json.loads(l) for l in open("ai-resource-definitions-en.corpus.jsonl")]
for c in chunks[:3]:
print(c.get("id"), "|", c.get("source"), "|", c["text"][:80])
# Vectorisez ensuite chaque c["text"] et gardez c["source"] pour citer la source dans la réponse
Communauté
Pas encore de commentaire. Partagez votre retour, il aidera les suivants.