Data Version Control | slacker news

1. dmpetrov ◴[19 Oct 24 20:39 UTC] No.41890616[source]▶

hi there! Maintainer and author here. Excited to see DVC on the front page!

Happy to answer any questions about DVC and our sister project DataChain https://github.com/iterative/datachain that does data versioning with a bit different assumptions: no file copy and built-in data transformations.

replies(3): >>41890932 #>>41896923 #>>41897005 #

2. ajoseps ◴[19 Oct 24 21:31 UTC] No.41890932[source]▶

>>41890616 (TP) #

if the data files are all just text files, what are the differences between DVC and using plain git?

replies(3): >>41891059 #>>41891080 #>>41893500 #

3. dmpetrov ◴[19 Oct 24 21:55 UTC] No.41891059[source]▶

>>41890932 #

In this cases, you need DVC if:

1. File are too large for Git and Git LFS.

2. You prefer using S3/GCS/Azure as a storage.

3. You need to track transformations/piplines on the file - clean up text file, train mode, etc.

Otherwise, vanilla Git may be sufficient.

4. miki123211 ◴[19 Oct 24 21:58 UTC] No.41891080[source]▶

>>41890932 #

DVC does a lot more than git.

It essentially makes sure that your results can reproducibly be generated from your original data. If any script or data file is changed, the parts of your pipeline that depend on it, possibly recursively, get re-run and the relevant results get updated automatically.

There's no chance of e.g. changing the structure of your original dataset slightly, forgetting to regenerate one of the intermediate models by accident, not noticing that the script to regenerate it doesn't work any more due to the new dataset structure, and then getting reminded a year later when moving to a new computer and trying to regen everything from scratch.

It's a lot like Unix make, but with the ability to keep track of different git branches and the data / intermediates they need, which saves you from needing to regen everything every time you make a new checkout, lets you easily exchange large datasets with teammates etc.

In theory, you could store everything in git, but then every time you made a small change to your scripts that e.g. changed the way some model works and slightly adjusted a score for each of ten million rows, your diff would be 10m LOC, and all versions of that dataset would be stored in your repo, forever, making it unbelievably large.

replies(3): >>41891756 #>>41894861 #>>41895262 #

5. azinman2 ◴[19 Oct 24 23:52 UTC] No.41891756{3}[source]▶

>>41891080 #

So where do the adjusted 10M rows live instead? S3?

replies(1): >>41892535 #

6. thangngoc89 ◴[20 Oct 24 02:55 UTC] No.41892535{4}[source]▶

>>41891756 #

DVC support multiple remotes. S3 is one of them, there are also WebDAV, local FS, Google Drive, and a bunch of others. You could see the full list here [0]. Disclaimer: not affiliated with DVC in anyway, just a user.

[0] https://dvc.org/doc/user-guide/data-management/remote-storag...

7. agile-gift0262 ◴[20 Oct 24 07:03 UTC] No.41893500[source]▶

>>41890932 #

It's not just to manage file versioning. Yo can define a pipeline with different stages, the dependencies and outputs of each stage and DVC will figure out which stages need running depending on what dependencies have changed. Stages can also output metrics and plots, and DVC has utilities to expose, explore and compare those.

8. woodglyst ◴[20 Oct 24 12:20 UTC] No.41894861{3}[source]▶

>>41891080 #

This sounds a lot like the experimental project Jacquard [0] from Ink & Switch.

[0] https://www.inkandswitch.com/jacquard/notebook/

9. amelius ◴[20 Oct 24 13:32 UTC] No.41895262{3}[source]▶

>>41891080 #

Sounds like it is more a framework than a tool.

Not everybody wants a framework.

replies(2): >>41895874 #>>41896912 #

10. JadeNB ◴[20 Oct 24 15:20 UTC] No.41895874{4}[source]▶

>>41895262 #

> Sounds like it is more a framework than a tool.

> Not everybody wants a framework.

The second part of this comment seems strange to me. Surely nothing on Hacker News is shared with the expectation that it will be interesting, or useful, to everyone. Equally, surely there are some people on HN who will be interested in a framework, even if it might be too heavy for other people.

replies(1): >>41896274 #

11. amelius ◴[20 Oct 24 16:05 UTC] No.41896274{5}[source]▶

>>41895874 #

Just saying that what makes Git so appealing is that it does one thing well, and from this view DVC seems to be in an entirely different category.

12. stochastastic ◴[20 Oct 24 17:17 UTC] No.41896912{4}[source]▶

>>41895262 #

It doesn’t force you to use any of the extra functionality. My team has been using it just for the version control part for a couple years and it has worked great.

replies(1): >>41954372 #

13. johanneskanybal ◴[20 Oct 24 17:19 UTC] No.41896923[source]▶

>>41890616 (TP) #

Mostly consult as a data engineer not ML ops but I’m interested in some aspects of this. We have 10 years of parquet files from 300+ different kafka topic and we’re currently migrating to apache iceberg. We’ll back fill on a need only basis and it would be nice to track that with git. Would this be a good fit for that?

Another potential aspect would be tracking schema evolution in a nicer way than we currently do.

thx in advance, huge fan of anything-as-code and think it’s a great fit for data (20+ years in this area).

14. stochastastic ◴[20 Oct 24 17:30 UTC] No.41897005[source]▶

>>41890616 (TP) #

Thanks for making and sharing DVC! It’s been a big help.

Is there any support that would be helpful? I’ll look at the project page too.

replies(1): >>41897163 #

15. dmpetrov ◴[20 Oct 24 17:53 UTC] No.41897163[source]▶

>>41897005 #

Thank you!

Just shoot an email to support and mention HN. I’ll read and reply.

16. bach4ants ◴[26 Oct 24 12:33 UTC] No.41954372{5}[source]▶

>>41896912 #

Yep. I personally like DVC's pipeline implementation because it's lightweight and language-agnostic, but haven't gotten into using their experiment tracking features.