Corpus: AI resource definitions (English)
Definitions of every AI resource type on Contexte Tech, one chunk each with the guide URL: a small English corpus ready for RAG.
Suddivisione : 1 resource type = 1 chunk
Anteprima
| id | Testo | Fonte |
|---|---|---|
context | A context is a set of instructions and examples given to the AI model before the user's question. It steers the answer without retraining anything: this is in-context learning. | https://contextetech.com/en/docs/contexts |
prompt | A prompt is a reusable text template with variables in double braces, for example {{product}}. Fill in the variables, then send it to the model. | https://contextetech.com/en/docs/prompts |
dataset | A dataset is a set of examples in JSONL format (one JSON line per example) used to train or evaluate a model. | https://contextetech.com/en/docs/datasets |
lora | LoRA adapts an existing model to a task without fully retraining it: only small adapters are trained. A LoRA page gathers all the settings of that training. | https://contextetech.com/en/docs/lora |
tool | An MCP (Model Context Protocol) server gives an AI assistant new abilities: read files, query a database, call a service. The same server works with Claude, Cursor and other assistants. | https://contextetech.com/en/docs/mcp-servers |
model | A model is an AI model published by a verified Lab: fine-tune, merge or quantized version. The page describes the model and links to its download. | https://contextetech.com/en/docs/models |
agent | An agent is an AI assistant that chains steps to complete a task: it follows instructions, uses MCP servers and contexts, and respects guardrails. | https://contextetech.com/en/docs/agents |
skill | A skill is an instruction pack (SKILL.md file) that an assistant like Claude loads only when the task calls for it. It teaches a precise way of doing things. | https://contextetech.com/en/docs/skills |
eval | An evaluation is a set of test cases (input and expected answer) with a scoring method. It measures whether a model, prompt, context or agent does its job. | https://contextetech.com/en/docs/evals |
rag | A RAG (retrieval-augmented generation) pipeline lets a model answer from your documents: they are chunked and embedded, then for each question the useful excerpts are retrieved and given to the model. | https://contextetech.com/en/docs/rag |
harness | A harness is the environment that runs an agent: its instructions file (e.g. CLAUDE.md), the tools it may use, what is denied, hooks and connected MCP servers. | https://contextetech.com/en/docs/harness |
workflow | A workflow is a chain of steps where AI takes part: fetch data, have a model analyse it, then send the result elsewhere. It is built in a tool such as n8n, LangGraph, Make or Dify, then exported as JSON. | https://contextetech.com/en/docs/workflows |
guardrail | A guardrail is a set of rules that controls what goes into or out of an AI assistant: block personal data, detect prompt injection, refuse off-topic requests. | https://contextetech.com/en/docs/guardrails |
embedding | An embedding model turns a text into a list of numbers (a vector). Texts with similar meaning give close vectors: it is the basis of semantic search and RAG. | https://contextetech.com/en/docs/embeddings |
inference | An inference config is the set of settings needed to run a model on your own machine: the runtime (Ollama, vLLM, llama.cpp), quantization, context size and launch command. | https://contextetech.com/en/docs/inference |
stack | A stack is a set of AI services that work together and start with one command: for example a model, a vector database and a chat interface, described in a docker-compose file. | https://contextetech.com/en/docs/stacks |
redteam | Red teaming means deliberately attacking an AI assistant to find its weaknesses before users do: rule bypass, data leaks, prompt injection, promises it should not make. | https://contextetech.com/en/docs/red-teaming |
aiact | An AI Act file is a documentation template to fill in to comply with the EU AI Act: transparency notice, training data description, risk assessment or system register. | https://contextetech.com/en/docs/ai-act-files |
schema | An output schema is a JSON Schema describing exactly the shape of the expected answer: which fields, which types, which are required. Recent models can follow it (structured outputs). | https://contextetech.com/en/docs/schemas |
corpus | A corpus is a set of documents already split into chunks, each with its text and source. It is ready to embed for RAG, and every answer can cite where it comes from. | https://contextetech.com/en/docs/corpus |
Costo stimato
Token di input per chiamata : ≈ 3028
| Modello | Per chiamata | Per 1.000 chiamate |
|---|---|---|
| Mistral Large · Mistral AI | 0,00133 € | 1,33 € |
| DeepSeek Flash · DeepSeek | 0,00080 € | 0,80 € |
| DeepSeek V4 Pro · DeepSeek | 0,00352 € | 3,52 € |
| GPT-6 Luna · OpenAI | 0,00027 € | 0,27 € |
| GPT-6 Sol · OpenAI | 0,00533 € | 5,33 € |
| Claude Sonnet 5.5 · Anthropic | 0,00533 € | 5,33 € |
| Claude Opus 5.5 · Anthropic | 0,01 € | 10,67 € |
| Modello open source self-hosted (Ollama, vLLM) | 0 € di API (solo costo del server) | |
Solo costo di input, esclusa la risposta del modello. Prezzi pubblici standard dei fornitori (Mistral AI, DeepSeek, OpenAI, Anthropic) convertiti da USD in euro al cambio BCE del 29 settembre 2026. Stima: 1 token ≈ 3,6 caratteri.
Common to all 3 regions
Created by the author, no outside source.
Commercial use allowed · credit the author · changes allowed.
Text or data only: no access to files, network or commands.
Hosted in France (Scaleway, Paris). No transfer outside the EU.
European Union · 5/5
No personal data.
No training data: no summary required.
United States · 5/5
No personal data.
No training data: no documentation to publish.
China · 5/5
No personal data.
No training data; AI-generated content published in China must be labelled (2025).
Indicative summary as of 30/09/2026, not a legal certification. Method and sources →
Statistiche
Mi piace : 0 · Commenti : 0
Clicca su un contatore per mostrarlo o nasconderlo; passa sul grafico per i dettagli. Una visualizzazione per visitatore al giorno, bot esclusi; conteggio iniziato il 30 settembre 2026.
- Identificativo
- contextetech--ai-resource-definitions-en
- Tipo
- Corpus
- File
- ai-resource-definitions-en.corpus.jsonl
- Formato
- JSON Lines (application/jsonl)
- Dimensione
- 4852 byte (4,7 KB)
- Codifica
- UTF-8
- Token stimati
- ≈ 1348
- Lingua
- Inglese
- Licenza
- MIT
- Uso commerciale
- Consentito
- Dati personali
- Nessuno
- Versione
- 1.0
- Pubblicato il
- 1 ottobre 2026 alle ore 18:55
- Aggiornato il
- 1 ottobre 2026 alle ore 18:55
- Archiviazione
- nel database
- Utilizzi
- 0
- Mi piace
- 0
ai-resource-definitions-en.corpus.jsonl
{"id":"context","text":"A context is a set of instructions and examples given to the AI model before the user's question. It steers the answer without retraining anything: this is in-context learning.","source":"https://contextetech.com/en/docs/contexts"}
{"id":"prompt","text":"A prompt is a reusable text template with variables in double braces, for example {{product}}. Fill in the variables, then send it to the model.","source":"https://contextetech.com/en/docs/prompts"}
{"id":"dataset","text":"A dataset is a set of examples in JSONL format (one JSON line per example) used to train or evaluate a model.","source":"https://contextetech.com/en/docs/datasets"}
{"id":"lora","text":"LoRA adapts an existing model to a task without fully retraining it: only small adapters are trained. A LoRA page gathers all the settings of that training.","source":"https://contextetech.com/en/docs/lora"}
{"id":"tool","text":"An MCP (Model Context Protocol) server gives an AI assistant new abilities: read files, query a database, call a service. The same server works with Claude, Cursor and other assistants.","source":"https://contextetech.com/en/docs/mcp-servers"}
{"id":"model","text":"A model is an AI model published by a verified Lab: fine-tune, merge or quantized version. The page describes the model and links to its download.","source":"https://contextetech.com/en/docs/models"}
{"id":"agent","text":"An agent is an AI assistant that chains steps to complete a task: it follows instructions, uses MCP servers and contexts, and respects guardrails.","source":"https://contextetech.com/en/docs/agents"}
{"id":"skill","text":"A skill is an instruction pack (SKILL.md file) that an assistant like Claude loads only when the task calls for it. It teaches a precise way of doing things.","source":"https://contextetech.com/en/docs/skills"}
{"id":"eval","text":"An evaluation is a set of test cases (input and expected answer) with a scoring method. It measures whether a model, prompt, context or agent does its job.","source":"https://contextetech.com/en/docs/evals"}
{"id":"rag","text":"A RAG (retrieval-augmented generation) pipeline lets a model answer from your documents: they are chunked and embedded, then for each question the useful excerpts are retrieved and given to the model.","source":"https://contextetech.com/en/docs/rag"}
{"id":"harness","text":"A harness is the environment that runs an agent: its instructions file (e.g. CLAUDE.md), the tools it may use, what is denied, hooks and connected MCP servers.","source":"https://contextetech.com/en/docs/harness"}
{"id":"workflow","text":"A workflow is a chain of steps where AI takes part: fetch data, have a model analyse it, then send the result elsewhere. It is built in a tool such as n8n, LangGraph, Make or Dify, then exported as JSON.","source":"https://contextetech.com/en/docs/workflows"}
{"id":"guardrail","text":"A guardrail is a set of rules that controls what goes into or out of an AI assistant: block personal data, detect prompt injection, refuse off-topic requests.","source":"https://contextetech.com/en/docs/guardrails"}
{"id":"embedding","text":"An embedding model turns a text into a list of numbers (a vector). Texts with similar meaning give close vectors: it is the basis of semantic search and RAG.","source":"https://contextetech.com/en/docs/embeddings"}
{"id":"inference","text":"An inference config is the set of settings needed to run a model on your own machine: the runtime (Ollama, vLLM, llama.cpp), quantization, context size and launch command.","source":"https://contextetech.com/en/docs/inference"}
{"id":"stack","text":"A stack is a set of AI services that work together and start with one command: for example a model, a vector database and a chat interface, described in a docker-compose file.","source":"https://contextetech.com/en/docs/stacks"}
{"id":"redteam","text":"Red teaming means deliberately attacking an AI assistant to find its weaknesses before users do: rule bypass, data leaks, prompt injection, promises it should not make.","source":"https://contextetech.com/en/docs/red-teaming"}
{"id":"aiact","text":"An AI Act file is a documentation template to fill in to comply with the EU AI Act: transparency notice, training data description, risk assessment or system register.","source":"https://contextetech.com/en/docs/ai-act-files"}
{"id":"schema","text":"An output schema is a JSON Schema describing exactly the shape of the expected answer: which fields, which types, which are required. Recent models can follow it (structured outputs).","source":"https://contextetech.com/en/docs/schemas"}
{"id":"corpus","text":"A corpus is a set of documents already split into chunks, each with its text and source. It is ready to embed for RAG, and every answer can cite where it comes from.","source":"https://contextetech.com/en/docs/corpus"}
Caricare il corpus (Python)
import json
chunks = [json.loads(l) for l in open("ai-resource-definitions-en.corpus.jsonl")]
for c in chunks[:3]:
print(c.get("id"), "|", c.get("source"), "|", c["text"][:80])
# Vectorisez ensuite chaque c["text"] et gardez c["source"] pour citer la source dans la réponse
Community
Ancora nessun commento. Condividi la tua esperienza, aiuterà chi viene dopo.