Dataset LLMs & Texto

HuggingFaceFW/fineweb

Dataset com mais de 1 trilhão de exemplos — 598.8 mil downloads no Hugging Face. 🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer What is it? The 🍷 FineWeb dataset consists of more than 18.

Hugging Face · Datasets ·HuggingFaceFW · ·↓ 598811 ·♥ 2992

O dataset HuggingFaceFW/fineweb está entre os destaques do Hugging Face — dados que alimentam o treinamento e a avaliação dos modelos do momento.

Ficha do dataset

  • Tamanho: mais de 1 trilhão de exemplos
  • Tarefas: geração de texto
  • Idiomas: inglês
  • Licença: ODC-BY
  • Downloads: 598.8 mil · Curtidas: 3.0 mil

Sobre o dataset

🍷 FineWeb 15 trillion tokens of the finest data the 🌐 web has to offer What is it? The 🍷 FineWeb dataset consists of more than 18.

Como carregar

Use a biblioteca datasets do Hugging Face:

pip install -U datasets

Como é um dataset grande, vale carregar em modo streaming (sem baixar tudo):

from datasets import load_dataset

ds = load_dataset("HuggingFaceFW/fineweb", split="train", streaming=True)
for exemplo in ds.take(3):
    print(exemplo)

Tags

text-generation

Explorar o dataset no Hugging Face →

compartilhar: