Corpus: AI resource definitions (English)
Definitions of every AI resource type on Contexte Tech, one chunk each with the guide URL: a small English corpus ready for RAG.
Fragmentación : 1 resource type = 1 chunk
Vista previa
| id | Texto | Fuente |
|---|---|---|
context | A context is a set of instructions and examples given to the AI model before the user's question. It steers the answer without retraining anything: this is in-context learning. | https://contextetech.com/en/docs/contexts |
prompt | A prompt is a reusable text template with variables in double braces, for example {{product}}. Fill in the variables, then send it to the model. | https://contextetech.com/en/docs/prompts |
dataset | A dataset is a set of examples in JSONL format (one JSON line per example) used to train or evaluate a model. | https://contextetech.com/en/docs/datasets |
lora | LoRA adapts an existing model to a task without fully retraining it: only small adapters are trained. A LoRA page gathers all the settings of that training. | https://contextetech.com/en/docs/lora |
tool | An MCP (Model Context Protocol) server gives an AI assistant new abilities: read files, query a database, call a service. The same server works with Claude, Cursor and other assistants. | https://contextetech.com/en/docs/mcp-servers |
model | A model is an AI model published by a verified Lab: fine-tune, merge or quantized version. The page describes the model and links to its download. | https://contextetech.com/en/docs/models |
agent | An agent is an AI assistant that chains steps to complete a task: it follows instructions, uses MCP servers and contexts, and respects guardrails. | https://contextetech.com/en/docs/agents |
skill | A skill is an instruction pack (SKILL.md file) that an assistant like Claude loads only when the task calls for it. It teaches a precise way of doing things. | https://contextetech.com/en/docs/skills |
eval | An evaluation is a set of test cases (input and expected answer) with a scoring method. It measures whether a model, prompt, context or agent does its job. | https://contextetech.com/en/docs/evals |
rag | A RAG (retrieval-augmented generation) pipeline lets a model answer from your documents: they are chunked and embedded, then for each question the useful excerpts are retrieved and given to the model. | https://contextetech.com/en/docs/rag |
harness | A harness is the environment that runs an agent: its instructions file (e.g. CLAUDE.md), the tools it may use, what is denied, hooks and connected MCP servers. | https://contextetech.com/en/docs/harness |
workflow | A workflow is a chain of steps where AI takes part: fetch data, have a model analyse it, then send the result elsewhere. It is built in a tool such as n8n, LangGraph, Make or Dify, then exported as JSON. | https://contextetech.com/en/docs/workflows |
guardrail | A guardrail is a set of rules that controls what goes into or out of an AI assistant: block personal data, detect prompt injection, refuse off-topic requests. | https://contextetech.com/en/docs/guardrails |
embedding | An embedding model turns a text into a list of numbers (a vector). Texts with similar meaning give close vectors: it is the basis of semantic search and RAG. | https://contextetech.com/en/docs/embeddings |
inference | An inference config is the set of settings needed to run a model on your own machine: the runtime (Ollama, vLLM, llama.cpp), quantization, context size and launch command. | https://contextetech.com/en/docs/inference |
stack | A stack is a set of AI services that work together and start with one command: for example a model, a vector database and a chat interface, described in a docker-compose file. | https://contextetech.com/en/docs/stacks |
redteam | Red teaming means deliberately attacking an AI assistant to find its weaknesses before users do: rule bypass, data leaks, prompt injection, promises it should not make. | https://contextetech.com/en/docs/red-teaming |
aiact | An AI Act file is a documentation template to fill in to comply with the EU AI Act: transparency notice, training data description, risk assessment or system register. | https://contextetech.com/en/docs/ai-act-files |
schema | An output schema is a JSON Schema describing exactly the shape of the expected answer: which fields, which types, which are required. Recent models can follow it (structured outputs). | https://contextetech.com/en/docs/schemas |
corpus | A corpus is a set of documents already split into chunks, each with its text and source. It is ready to embed for RAG, and every answer can cite where it comes from. | https://contextetech.com/en/docs/corpus |
Coste estimado
Tokens de entrada por llamada : ≈ 3028
| Modelo | Por llamada | Por 1.000 llamadas |
|---|---|---|
| Mistral Large · Mistral AI | 0,00133 € | 1,33 € |
| DeepSeek Flash · DeepSeek | 0,00080 € | 0,80 € |
| DeepSeek V4 Pro · DeepSeek | 0,00352 € | 3,52 € |
| GPT-6 Luna · OpenAI | 0,00027 € | 0,27 € |
| GPT-6 Sol · OpenAI | 0,00533 € | 5,33 € |
| Claude Sonnet 5.5 · Anthropic | 0,00533 € | 5,33 € |
| Claude Opus 5.5 · Anthropic | 0,01 € | 10,67 € |
| Modelo de código abierto autoalojado (Ollama, vLLM) | 0 € de API (solo el coste del servidor) | |
Solo coste de entrada, sin la respuesta del modelo. Precios públicos estándar de los proveedores (Mistral AI, DeepSeek, OpenAI, Anthropic) convertidos de USD a euros al tipo del BCE del 29 de septiembre de 2026. Estimación: 1 token ≈ 3,6 caracteres.
Common to all 3 regions
Created by the author, no outside source.
Commercial use allowed · credit the author · changes allowed.
Text or data only: no access to files, network or commands.
Hosted in France (Scaleway, Paris). No transfer outside the EU.
European Union · 5/5
No personal data.
No training data: no summary required.
United States · 5/5
No personal data.
No training data: no documentation to publish.
China · 5/5
No personal data.
No training data; AI-generated content published in China must be labelled (2025).
Indicative summary as of 30/09/2026, not a legal certification. Method and sources →
Estadísticas
Me gusta : 0 · Comentarios : 0
Haz clic en un contador para mostrarlo u ocultarlo; pasa el ratón por el gráfico para ver el detalle. Una visita por visitante y día, sin robots; recuento iniciado el 30 de septiembre de 2026.
- Identificador
- contextetech--ai-resource-definitions-en
- Tipo
- Corpus
- Archivo
- ai-resource-definitions-en.corpus.jsonl
- Formato
- JSON Lines (application/jsonl)
- Tamaño
- 4852 bytes (4,7 KB)
- Codificación
- UTF-8
- Tokens estimados
- ≈ 1348
- Idioma
- Inglés
- Licencia
- MIT
- Uso comercial
- Permitido
- Datos personales
- Ninguno
- Versión
- 1.0
- Publicado el
- 1 de octubre de 2026 a las 18:55
- Actualizado el
- 1 de octubre de 2026 a las 18:55
- Almacenamiento
- en base de datos
- Usos
- 0
- Me gusta
- 0
ai-resource-definitions-en.corpus.jsonl
{"id":"context","text":"A context is a set of instructions and examples given to the AI model before the user's question. It steers the answer without retraining anything: this is in-context learning.","source":"https://contextetech.com/en/docs/contexts"}
{"id":"prompt","text":"A prompt is a reusable text template with variables in double braces, for example {{product}}. Fill in the variables, then send it to the model.","source":"https://contextetech.com/en/docs/prompts"}
{"id":"dataset","text":"A dataset is a set of examples in JSONL format (one JSON line per example) used to train or evaluate a model.","source":"https://contextetech.com/en/docs/datasets"}
{"id":"lora","text":"LoRA adapts an existing model to a task without fully retraining it: only small adapters are trained. A LoRA page gathers all the settings of that training.","source":"https://contextetech.com/en/docs/lora"}
{"id":"tool","text":"An MCP (Model Context Protocol) server gives an AI assistant new abilities: read files, query a database, call a service. The same server works with Claude, Cursor and other assistants.","source":"https://contextetech.com/en/docs/mcp-servers"}
{"id":"model","text":"A model is an AI model published by a verified Lab: fine-tune, merge or quantized version. The page describes the model and links to its download.","source":"https://contextetech.com/en/docs/models"}
{"id":"agent","text":"An agent is an AI assistant that chains steps to complete a task: it follows instructions, uses MCP servers and contexts, and respects guardrails.","source":"https://contextetech.com/en/docs/agents"}
{"id":"skill","text":"A skill is an instruction pack (SKILL.md file) that an assistant like Claude loads only when the task calls for it. It teaches a precise way of doing things.","source":"https://contextetech.com/en/docs/skills"}
{"id":"eval","text":"An evaluation is a set of test cases (input and expected answer) with a scoring method. It measures whether a model, prompt, context or agent does its job.","source":"https://contextetech.com/en/docs/evals"}
{"id":"rag","text":"A RAG (retrieval-augmented generation) pipeline lets a model answer from your documents: they are chunked and embedded, then for each question the useful excerpts are retrieved and given to the model.","source":"https://contextetech.com/en/docs/rag"}
{"id":"harness","text":"A harness is the environment that runs an agent: its instructions file (e.g. CLAUDE.md), the tools it may use, what is denied, hooks and connected MCP servers.","source":"https://contextetech.com/en/docs/harness"}
{"id":"workflow","text":"A workflow is a chain of steps where AI takes part: fetch data, have a model analyse it, then send the result elsewhere. It is built in a tool such as n8n, LangGraph, Make or Dify, then exported as JSON.","source":"https://contextetech.com/en/docs/workflows"}
{"id":"guardrail","text":"A guardrail is a set of rules that controls what goes into or out of an AI assistant: block personal data, detect prompt injection, refuse off-topic requests.","source":"https://contextetech.com/en/docs/guardrails"}
{"id":"embedding","text":"An embedding model turns a text into a list of numbers (a vector). Texts with similar meaning give close vectors: it is the basis of semantic search and RAG.","source":"https://contextetech.com/en/docs/embeddings"}
{"id":"inference","text":"An inference config is the set of settings needed to run a model on your own machine: the runtime (Ollama, vLLM, llama.cpp), quantization, context size and launch command.","source":"https://contextetech.com/en/docs/inference"}
{"id":"stack","text":"A stack is a set of AI services that work together and start with one command: for example a model, a vector database and a chat interface, described in a docker-compose file.","source":"https://contextetech.com/en/docs/stacks"}
{"id":"redteam","text":"Red teaming means deliberately attacking an AI assistant to find its weaknesses before users do: rule bypass, data leaks, prompt injection, promises it should not make.","source":"https://contextetech.com/en/docs/red-teaming"}
{"id":"aiact","text":"An AI Act file is a documentation template to fill in to comply with the EU AI Act: transparency notice, training data description, risk assessment or system register.","source":"https://contextetech.com/en/docs/ai-act-files"}
{"id":"schema","text":"An output schema is a JSON Schema describing exactly the shape of the expected answer: which fields, which types, which are required. Recent models can follow it (structured outputs).","source":"https://contextetech.com/en/docs/schemas"}
{"id":"corpus","text":"A corpus is a set of documents already split into chunks, each with its text and source. It is ready to embed for RAG, and every answer can cite where it comes from.","source":"https://contextetech.com/en/docs/corpus"}
Cargar el corpus (Python)
import json
chunks = [json.loads(l) for l in open("ai-resource-definitions-en.corpus.jsonl")]
for c in chunks[:3]:
print(c.get("id"), "|", c.get("source"), "|", c["text"][:80])
# Vectorisez ensuite chaque c["text"] et gardez c["source"] pour citer la source dans la réponse
Comunidad
Aún no hay comentarios. Comparte tu opinión, ayudará a los siguientes.