Any data scientists out there? What's your go to programming language and tools for your work?

rutrum@lm.paradisus.day · 4 months ago

Any data scientists out there? What's your go to programming language and tools for your work?

Kache@lemm.ee · 4 months ago

What kind of query optimization can it for scanning data that’s already in memory?

rutrum@lm.paradisus.day · 4 months ago

A big feature of polars is only loading applicable data from disk. But during exporatory data analysis (EDA) you often have the whole dataset in memory. In this case, filters wont help much there. Polars has a good page in their docs about all the possible optimizations it is capable of. https://docs.pola.rs/user-guide/lazy/optimizations/

One I see off the top is projection pushdown, which only selects relevant columns for a final transformations. In pandas, if you perform a group by with aggregation, then only look at a few columns, you still perform aggregation across all the data. In polars lazy API, you would define the entire process upfront, and it would know not to aggregate certain columns, for instance.

Kache@lemm.ee · edit-2 4 months ago

Hm, that’s kind of interesting

But my first reaction is that optimizations only at the “Python processing level” are going to be pretty limited since it’s not going to have metadata/statistics, and it’d depend heavily on the source data layout, e.g. CSV vs parquet