Dataset
LLMs & Texto
HuggingFaceFW/fineweb
Dataset com mais de 1 trilhão de exemplos — 598.8 mil downloads no Hugging Face. 🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer What is it? The 🍷 FineWeb dataset consists of more than 18.
Hugging Face · Datasets
·HuggingFaceFW
·
·↓ 598811
·♥ 2992
O dataset HuggingFaceFW/fineweb está entre os destaques do Hugging Face — dados que alimentam o treinamento e a avaliação dos modelos do momento.
Ficha do dataset
- Tamanho: mais de 1 trilhão de exemplos
- Tarefas: geração de texto
- Idiomas: inglês
- Licença: ODC-BY
- Downloads: 598.8 mil · Curtidas: 3.0 mil
Sobre o dataset
🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer What is it? The 🍷 FineWeb dataset consists of more than 18.
Como carregar
Use a biblioteca datasets do Hugging Face:
pip install -U datasets
Como é um dataset grande, vale carregar em modo streaming (sem baixar tudo):
from datasets import load_dataset
ds = load_dataset("HuggingFaceFW/fineweb", split="train", streaming=True)
for exemplo in ds.take(3):
print(exemplo)Tags
text-generation