(ethz.ch)

524 points andy99 | 1 comments | 11 Jul 25 18:45 UTC | HN request time: 0.369s | source

Show context

WeirderScience ◴[11 Jul 25 20:05 UTC] No.44536327[source]▶

The open training data is a huge differentiator. Is this the first truly open dataset of this scale? Prior efforts like The Pile were valuable, but had limitations. Curious to see how reproducible the training is.

replies(2): >>44536400 #>>44537249 #

layer8 ◴[11 Jul 25 20:16 UTC] No.44536400[source]▶

>>44536327 #

> The model will be fully open: source code and weights will be publicly available, and the training data will be transparent and reproducible

This leads me to believe that the training data won’t be made publicly available in full, but merely be “reproducible”. This might mean that they’ll provide references like a list of URLs of the pages they trained on, but not their contents.

replies(3): >>44536448 #>>44536623 #>>44536818 #

TobTobXX ◴[11 Jul 25 21:11 UTC] No.44536818[source]▶

>>44536400 #

Well, when the actual content is 100s of terabytes big, providing URLs may be more practical for them and for others.

replies(1): >>44537342 #

1. layer8 ◴[11 Jul 25 22:19 UTC] No.44537342[source]▶

>>44536818 #

The difference between content they are allowed to train on vs. being allowed to distribute copies of is likely at least as relevant.

↑

ETH Zurich and EPFL to release a LLM developed on public infrastructure