Embeddings multilingual-e5-large für ein deutsches RAG
multilingual-e5-large für ein RAG auf Deutsch: 1024 Dimensionen, 512 Tokens, Präfixe query:/passage: erforderlich, Kosinus-Distanz. MIT-Lizenz.
| Modello | intfloat/multilingual-e5-large |
|---|---|
| Dimensioni | 1024 |
| Token max | 512 |
| Distanza | cosine |
| Lingue | de, en, fr, es, it, nl |
| Prefissi | query: / passage: |
| Dimensione dei chunk | 350 |
Note
Die Präfixe „query: “ und „passage: “ sind Pflicht: ohne sie wird die Suche deutlich schlechter. Höchstens 512 Tokens pro Text: Dokumente in Abschnitte von etwa 350 Wörtern teilen. Vektoren normalisieren und Kosinus-Distanz verwenden. Von Microsoft (intfloat) unter MIT-Lizenz veröffentlicht.
Costo stimato
Token di input per chiamata : ≈ 249
| Modello | Per chiamata | Per 1.000 chiamate |
|---|---|---|
| Mistral Large · Mistral AI | 0,00011 € | 0,11 € |
| DeepSeek Flash · DeepSeek | 0,00007 € | 0,07 € |
| DeepSeek V4 Pro · DeepSeek | 0,00029 € | 0,29 € |
| GPT-6 Luna · OpenAI | 0,00002 € | 0,02 € |
| GPT-6 Sol · OpenAI | 0,00044 € | 0,44 € |
| Claude Sonnet 5.5 · Anthropic | 0,00044 € | 0,44 € |
| Claude Opus 5.5 · Anthropic | 0,00088 € | 0,88 € |
| Modello open source self-hosted (Ollama, vLLM) | 0 € di API (solo costo del server) | |
Solo costo di input, esclusa la risposta del modello. Prezzi pubblici standard dei fornitori (Mistral AI, DeepSeek, OpenAI, Anthropic) convertiti da USD in euro al cambio BCE del 29 settembre 2026. Stima: 1 token ≈ 3,6 caratteri.
Common to all 3 regions
Sources under a compatible licence.
Sources : Modell intfloat/multilingual-e5-large, MIT-Lizenz (Hugging-Face-Modellkarte)
Commercial use allowed · credit the author · changes allowed.
Text or data only: no access to files, network or commands.
Hosted in France (Scaleway, Paris). No transfer outside the EU.
European Union · 5/5
No personal data.
No training data: no summary required.
United States · 5/5
No personal data.
No training data: no documentation to publish.
China · 5/5
No personal data.
No training data; AI-generated content published in China must be labelled (2025).
Indicative summary as of 30/09/2026, not a legal certification. Method and sources →
Statistiche
Mi piace : 0 · Commenti : 0
Clicca su un contatore per mostrarlo o nasconderlo; passa sul grafico per i dettagli. Una visualizzazione per visitatore al giorno, bot esclusi; conteggio iniziato il 30 settembre 2026.
- Identificativo
- contextetech--multilingual-e5-large-rag-deutsch
- Tipo
- Embedding
- File
- multilingual-e5-large-rag-deutsch.embedding.json
- Formato
- JSON (application/json)
- Dimensione
- 683 byte
- Codifica
- UTF-8
- Token stimati
- ≈ 189
- Lingua
- Tedesco
- Licenza
- MIT
- Uso commerciale
- Consentito
- Dati personali
- Nessuno
- Versione
- 1.0
- Pubblicato il
- 3 ottobre 2026 alle ore 05:20
- Aggiornato il
- 3 ottobre 2026 alle ore 05:20
- Archiviazione
- nel database
- Utilizzi
- 0
- Mi piace
- 0
multilingual-e5-large-rag-deutsch.embedding.json
{
"format": "contextetech/embedding/v1",
"id": "contextetech/multilingual-e5-large-rag-deutsch",
"description": "multilingual-e5-large für ein RAG auf Deutsch: 1024 Dimensionen, 512 Tokens, Präfixe query:/passage: erforderlich, Kosinus-Distanz. MIT-Lizenz.",
"tags": [
"rag",
"mehrsprachig",
"semantische-suche"
],
"license": "MIT",
"language": "de",
"model": "intfloat/multilingual-e5-large",
"dimensions": 1024,
"maxTokens": 512,
"distance": "cosine",
"languages": [
"de",
"en",
"fr",
"es",
"it",
"nl"
],
"prefixes": {
"query": "query: ",
"passage": "passage: "
},
"chunkSize": 350,
"benchmarks": []
}Vettorizzare (Python)
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("intfloat/multilingual-e5-large")
q = model.encode(["query: il tuo input qui"], normalize_embeddings=True)
docs = model.encode(["passage: …"], normalize_embeddings=True)
print(q.shape) # (1, 1024)
print(q @ docs.T) # similarité
Community
Ancora nessun commento. Condividi la tua esperienza, aiuterà chi viene dopo.