Embeddings multilingual-e5-large per un RAG in italiano
multilingual-e5-large per un RAG in italiano: 1024 dimensioni, 512 token, prefissi query:/passage: obbligatori, distanza coseno. Licenza MIT.
| Modell | intfloat/multilingual-e5-large |
|---|---|
| Dimensionen | 1024 |
| Max. Tokens | 512 |
| Distanz | cosine |
| Sprachen | it, en, fr, de, es, pt |
| Präfixe | query: / passage: |
| Chunk-Größe | 350 |
Notizen
I prefissi «query: » e «passage: » sono obbligatori: senza, la ricerca peggiora molto. Al massimo 512 token per testo: dividi i documenti in frammenti di circa 350 parole. Normalizza i vettori e usa la distanza coseno. Modello pubblicato da Microsoft (intfloat) con licenza MIT.
Geschätzte Kosten
Eingabe-Tokens pro Aufruf : ≈ 244
| Modell | Pro Aufruf | Pro 1.000 Aufrufe |
|---|---|---|
| Mistral Large · Mistral AI | 0,00011 € | 0,11 € |
| DeepSeek Flash · DeepSeek | 0,00006 € | 0,06 € |
| DeepSeek V4 Pro · DeepSeek | 0,00028 € | 0,28 € |
| GPT-6 Luna · OpenAI | 0,00002 € | 0,02 € |
| GPT-6 Sol · OpenAI | 0,00043 € | 0,43 € |
| Claude Sonnet 5.5 · Anthropic | 0,00043 € | 0,43 € |
| Claude Opus 5.5 · Anthropic | 0,00086 € | 0,86 € |
| Selbst gehostetes Open-Source-Modell (Ollama, vLLM) | 0 € API-Kosten (nur Serverkosten) | |
Nur Eingabekosten, ohne die Antwort des Modells. Öffentliche Standardpreise der Anbieter (Mistral AI, DeepSeek, OpenAI, Anthropic) von USD in Euro umgerechnet zum EZB-Kurs vom 29. September 2026. Schätzung: 1 Token ≈ 3,6 Zeichen.
Common to all 3 regions
Sources under a compatible licence.
Sources : Modello intfloat/multilingual-e5-large, licenza MIT (scheda Hugging Face)
Commercial use allowed · credit the author · changes allowed.
Text or data only: no access to files, network or commands.
Hosted in France (Scaleway, Paris). No transfer outside the EU.
European Union · 5/5
No personal data.
No training data: no summary required.
United States · 5/5
No personal data.
No training data: no documentation to publish.
China · 5/5
No personal data.
No training data; AI-generated content published in China must be labelled (2025).
Indicative summary as of 30/09/2026, not a legal certification. Method and sources →
Statistiken
Gefällt mir : 0 · Kommentare : 0
Klicken Sie auf einen Zähler, um ihn ein- oder auszublenden; fahren Sie über das Diagramm für Details. Ein Aufruf pro Besucher und Tag, ohne Bots; Zählung seit dem 30. September 2026.
- Kennung
- contextetech--multilingual-e5-large-rag-italiano
- Typ
- Embeddings
- Datei
- multilingual-e5-large-rag-italiano.embedding.json
- Format
- JSON (application/json)
- Größe
- 679 Bytes
- Kodierung
- UTF-8
- Geschätzte Tokens
- ≈ 189
- Sprache
- Italienisch
- Lizenz
- MIT
- Kommerzielle Nutzung
- Erlaubt
- Personenbezogene Daten
- Keine
- Version
- 1.0
- Veröffentlicht am
- 3. Oktober 2026 um 05:20
- Aktualisiert am
- 3. Oktober 2026 um 05:20
- Speicherung
- in der Datenbank
- Nutzungen
- 0
- Gefällt mir
- 0
multilingual-e5-large-rag-italiano.embedding.json
{
"format": "contextetech/embedding/v1",
"id": "contextetech/multilingual-e5-large-rag-italiano",
"description": "multilingual-e5-large per un RAG in italiano: 1024 dimensioni, 512 token, prefissi query:/passage: obbligatori, distanza coseno. Licenza MIT.",
"tags": [
"rag",
"multilingue",
"ricerca-semantica"
],
"license": "MIT",
"language": "it",
"model": "intfloat/multilingual-e5-large",
"dimensions": 1024,
"maxTokens": 512,
"distance": "cosine",
"languages": [
"it",
"en",
"fr",
"de",
"es",
"pt"
],
"prefixes": {
"query": "query: ",
"passage": "passage: "
},
"chunkSize": 350,
"benchmarks": []
}Vektorisieren (Python)
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("intfloat/multilingual-e5-large")
q = model.encode(["query: deine Eingabe hier"], normalize_embeddings=True)
docs = model.encode(["passage: …"], normalize_embeddings=True)
print(q.shape) # (1, 1024)
print(q @ docs.T) # similarité
Community
Noch keine Kommentare. Teile deine Erfahrung, sie hilft den Nächsten.