| FazBrowse GitHub Viewer | Trending | | Home |
| Tools: [Download Repo ZIP] [Original HTTPS Page] |
❌ 2 Tests Failed:
To view more test analytics, go to the Test Analytics Dashboard |
Sorry, something went wrong.
|
Reading lots of files can definitely cause headaches. However, auto-sharding with dask made my jobs extremely slow (I typically work with datasets of ~22 million observations by 1000 variables). |
Sorry, something went wrong.
|
But write-once (slow) and read-many (fast) seems a worthy tradeoff? |
Sorry, something went wrong.
It depends. In my case, sharding was causing my jobs to take an inordinate amount of time and my subsequent steps are CPU and not I/O bound. |
Sorry, something went wrong.
|
If you can't write the array, then you can't use the array afterwards. That seems strong enough. |
Sorry, something went wrong.
| Back | FazBrowse Home | New Git URL |
@joshua-gould I just noticed this also disables sharing for on-disk concatenation, which despite being a feature I dislike, does exist. I'm curious, before this gets merged into a release line, what your experience has been with storing/reading the data. Generally, lots of large files create headaches. Are you not seeing this?