Corpus: AI resource definitions (English)
Definitions of every AI resource type on Contexte Tech, one chunk each with the guide URL: a small English corpus ready for RAG.
Zerlegung : 1 resource type = 1 chunk
Vorschau
| id | Text | Quelle |
|---|---|---|
context | A context is a set of instructions and examples given to the AI model before the user's question. It steers the answer without retraining anything: this is in-context learning. | https://contextetech.com/en/docs/contexts |
prompt | A prompt is a reusable text template with variables in double braces, for example {{product}}. Fill in the variables, then send it to the model. | https://contextetech.com/en/docs/prompts |
dataset | A dataset is a set of examples in JSONL format (one JSON line per example) used to train or evaluate a model. | https://contextetech.com/en/docs/datasets |
lora | LoRA adapts an existing model to a task without fully retraining it: only small adapters are trained. A LoRA page gathers all the settings of that training. | https://contextetech.com/en/docs/lora |
tool | An MCP (Model Context Protocol) server gives an AI assistant new abilities: read files, query a database, call a service. The same server works with Claude, Cursor and other assistants. | https://contextetech.com/en/docs/mcp-servers |
model | A model is an AI model published by a verified Lab: fine-tune, merge or quantized version. The page describes the model and links to its download. | https://contextetech.com/en/docs/models |
agent | An agent is an AI assistant that chains steps to complete a task: it follows instructions, uses MCP servers and contexts, and respects guardrails. | https://contextetech.com/en/docs/agents |
skill | A skill is an instruction pack (SKILL.md file) that an assistant like Claude loads only when the task calls for it. It teaches a precise way of doing things. | https://contextetech.com/en/docs/skills |
eval | An evaluation is a set of test cases (input and expected answer) with a scoring method. It measures whether a model, prompt, context or agent does its job. | https://contextetech.com/en/docs/evals |
rag | A RAG (retrieval-augmented generation) pipeline lets a model answer from your documents: they are chunked and embedded, then for each question the useful excerpts are retrieved and given to the model. | https://contextetech.com/en/docs/rag |
harness | A harness is the environment that runs an agent: its instructions file (e.g. CLAUDE.md), the tools it may use, what is denied, hooks and connected MCP servers. | https://contextetech.com/en/docs/harness |
workflow | A workflow is a chain of steps where AI takes part: fetch data, have a model analyse it, then send the result elsewhere. It is built in a tool such as n8n, LangGraph, Make or Dify, then exported as JSON. | https://contextetech.com/en/docs/workflows |
guardrail | A guardrail is a set of rules that controls what goes into or out of an AI assistant: block personal data, detect prompt injection, refuse off-topic requests. | https://contextetech.com/en/docs/guardrails |
embedding | An embedding model turns a text into a list of numbers (a vector). Texts with similar meaning give close vectors: it is the basis of semantic search and RAG. | https://contextetech.com/en/docs/embeddings |
inference | An inference config is the set of settings needed to run a model on your own machine: the runtime (Ollama, vLLM, llama.cpp), quantization, context size and launch command. | https://contextetech.com/en/docs/inference |
stack | A stack is a set of AI services that work together and start with one command: for example a model, a vector database and a chat interface, described in a docker-compose file. | https://contextetech.com/en/docs/stacks |
redteam | Red teaming means deliberately attacking an AI assistant to find its weaknesses before users do: rule bypass, data leaks, prompt injection, promises it should not make. | https://contextetech.com/en/docs/red-teaming |
aiact | An AI Act file is a documentation template to fill in to comply with the EU AI Act: transparency notice, training data description, risk assessment or system register. | https://contextetech.com/en/docs/ai-act-files |
schema | An output schema is a JSON Schema describing exactly the shape of the expected answer: which fields, which types, which are required. Recent models can follow it (structured outputs). | https://contextetech.com/en/docs/schemas |
corpus | A corpus is a set of documents already split into chunks, each with its text and source. It is ready to embed for RAG, and every answer can cite where it comes from. | https://contextetech.com/en/docs/corpus |
Geschätzte Kosten
Eingabe-Tokens pro Aufruf : ≈ 3.028
| Modell | Pro Aufruf | Pro 1.000 Aufrufe |
|---|---|---|
| Mistral Large · Mistral AI | 0,00133 € | 1,33 € |
| DeepSeek Flash · DeepSeek | 0,00080 € | 0,80 € |
| DeepSeek V4 Pro · DeepSeek | 0,00352 € | 3,52 € |
| GPT-6 Luna · OpenAI | 0,00027 € | 0,27 € |
| GPT-6 Sol · OpenAI | 0,00533 € | 5,33 € |
| Claude Sonnet 5.5 · Anthropic | 0,00533 € | 5,33 € |
| Claude Opus 5.5 · Anthropic | 0,01 € | 10,67 € |
| Selbst gehostetes Open-Source-Modell (Ollama, vLLM) | 0 € API-Kosten (nur Serverkosten) | |
Nur Eingabekosten, ohne die Antwort des Modells. Öffentliche Standardpreise der Anbieter (Mistral AI, DeepSeek, OpenAI, Anthropic) von USD in Euro umgerechnet zum EZB-Kurs vom 29. September 2026. Schätzung: 1 Token ≈ 3,6 Zeichen.
Common to all 3 regions
Created by the author, no outside source.
Commercial use allowed · credit the author · changes allowed.
Text or data only: no access to files, network or commands.
Hosted in France (Scaleway, Paris). No transfer outside the EU.
European Union · 5/5
No personal data.
No training data: no summary required.
United States · 5/5
No personal data.
No training data: no documentation to publish.
China · 5/5
No personal data.
No training data; AI-generated content published in China must be labelled (2025).
Indicative summary as of 30/09/2026, not a legal certification. Method and sources →
Statistiken
Gefällt mir : 0 · Kommentare : 0
Klicken Sie auf einen Zähler, um ihn ein- oder auszublenden; fahren Sie über das Diagramm für Details. Ein Aufruf pro Besucher und Tag, ohne Bots; Zählung seit dem 30. September 2026.
- Kennung
- contextetech--ai-resource-definitions-en
- Typ
- Korpora
- Datei
- ai-resource-definitions-en.corpus.jsonl
- Format
- JSON Lines (application/jsonl)
- Größe
- 4.852 Bytes (4,7 KB)
- Kodierung
- UTF-8
- Geschätzte Tokens
- ≈ 1.348
- Sprache
- Englisch
- Lizenz
- MIT
- Kommerzielle Nutzung
- Erlaubt
- Personenbezogene Daten
- Keine
- Version
- 1.0
- Veröffentlicht am
- 1. Oktober 2026 um 18:55
- Aktualisiert am
- 1. Oktober 2026 um 18:55
- Speicherung
- in der Datenbank
- Nutzungen
- 0
- Gefällt mir
- 0
ai-resource-definitions-en.corpus.jsonl
{"id":"context","text":"A context is a set of instructions and examples given to the AI model before the user's question. It steers the answer without retraining anything: this is in-context learning.","source":"https://contextetech.com/en/docs/contexts"}
{"id":"prompt","text":"A prompt is a reusable text template with variables in double braces, for example {{product}}. Fill in the variables, then send it to the model.","source":"https://contextetech.com/en/docs/prompts"}
{"id":"dataset","text":"A dataset is a set of examples in JSONL format (one JSON line per example) used to train or evaluate a model.","source":"https://contextetech.com/en/docs/datasets"}
{"id":"lora","text":"LoRA adapts an existing model to a task without fully retraining it: only small adapters are trained. A LoRA page gathers all the settings of that training.","source":"https://contextetech.com/en/docs/lora"}
{"id":"tool","text":"An MCP (Model Context Protocol) server gives an AI assistant new abilities: read files, query a database, call a service. The same server works with Claude, Cursor and other assistants.","source":"https://contextetech.com/en/docs/mcp-servers"}
{"id":"model","text":"A model is an AI model published by a verified Lab: fine-tune, merge or quantized version. The page describes the model and links to its download.","source":"https://contextetech.com/en/docs/models"}
{"id":"agent","text":"An agent is an AI assistant that chains steps to complete a task: it follows instructions, uses MCP servers and contexts, and respects guardrails.","source":"https://contextetech.com/en/docs/agents"}
{"id":"skill","text":"A skill is an instruction pack (SKILL.md file) that an assistant like Claude loads only when the task calls for it. It teaches a precise way of doing things.","source":"https://contextetech.com/en/docs/skills"}
{"id":"eval","text":"An evaluation is a set of test cases (input and expected answer) with a scoring method. It measures whether a model, prompt, context or agent does its job.","source":"https://contextetech.com/en/docs/evals"}
{"id":"rag","text":"A RAG (retrieval-augmented generation) pipeline lets a model answer from your documents: they are chunked and embedded, then for each question the useful excerpts are retrieved and given to the model.","source":"https://contextetech.com/en/docs/rag"}
{"id":"harness","text":"A harness is the environment that runs an agent: its instructions file (e.g. CLAUDE.md), the tools it may use, what is denied, hooks and connected MCP servers.","source":"https://contextetech.com/en/docs/harness"}
{"id":"workflow","text":"A workflow is a chain of steps where AI takes part: fetch data, have a model analyse it, then send the result elsewhere. It is built in a tool such as n8n, LangGraph, Make or Dify, then exported as JSON.","source":"https://contextetech.com/en/docs/workflows"}
{"id":"guardrail","text":"A guardrail is a set of rules that controls what goes into or out of an AI assistant: block personal data, detect prompt injection, refuse off-topic requests.","source":"https://contextetech.com/en/docs/guardrails"}
{"id":"embedding","text":"An embedding model turns a text into a list of numbers (a vector). Texts with similar meaning give close vectors: it is the basis of semantic search and RAG.","source":"https://contextetech.com/en/docs/embeddings"}
{"id":"inference","text":"An inference config is the set of settings needed to run a model on your own machine: the runtime (Ollama, vLLM, llama.cpp), quantization, context size and launch command.","source":"https://contextetech.com/en/docs/inference"}
{"id":"stack","text":"A stack is a set of AI services that work together and start with one command: for example a model, a vector database and a chat interface, described in a docker-compose file.","source":"https://contextetech.com/en/docs/stacks"}
{"id":"redteam","text":"Red teaming means deliberately attacking an AI assistant to find its weaknesses before users do: rule bypass, data leaks, prompt injection, promises it should not make.","source":"https://contextetech.com/en/docs/red-teaming"}
{"id":"aiact","text":"An AI Act file is a documentation template to fill in to comply with the EU AI Act: transparency notice, training data description, risk assessment or system register.","source":"https://contextetech.com/en/docs/ai-act-files"}
{"id":"schema","text":"An output schema is a JSON Schema describing exactly the shape of the expected answer: which fields, which types, which are required. Recent models can follow it (structured outputs).","source":"https://contextetech.com/en/docs/schemas"}
{"id":"corpus","text":"A corpus is a set of documents already split into chunks, each with its text and source. It is ready to embed for RAG, and every answer can cite where it comes from.","source":"https://contextetech.com/en/docs/corpus"}
Korpus laden (Python)
import json
chunks = [json.loads(l) for l in open("ai-resource-definitions-en.corpus.jsonl")]
for c in chunks[:3]:
print(c.get("id"), "|", c.get("source"), "|", c["text"][:80])
# Vectorisez ensuite chaque c["text"] et gardez c["source"] pour citer la source dans la réponse
Community
Noch keine Kommentare. Teile deine Erfahrung, sie hilft den Nächsten.