Analytics Insight published another 2026 tools list on September 27. Pardeep Sharma. The stack is the one you already have on a conference slide: Python and SQL at the base, NumPy and pandas for everyday frames, Polars and DuckDB as the faster locals, scikit-learn and XGBoost for classical ML, PyTorch and TensorFlow and Hugging Face for the rest, Spark when one machine is not enough, MLflow to watch models, Power BI or Tableau to make a chart a VP will accept.
DuckDB Lab, the day before, is not a list. It is DuckDB v1.5.x on Linux x86_64 with 16 GB of RAM, reading Hive-style Parquet partitions so WHERE year = 2026 AND month = 9 does not scan 33 months of files.
Those two URLs are the week. One is a poster. One is a file skip.
We already covered DuckLake putting a catalog in SQL and DuckDB as a Python analytics engine. This is the Parquet read path the list still treats as a bullet under “scale.”
The list is a seating chart
Sharma’s key takeaways are three buckets. Foundation: Python, SQL, NumPy, pandas. Modern ML: scikit-learn, XGBoost, PyTorch, TensorFlow, Hugging Face. Scale and communication: Polars, DuckDB, Spark, MLflow, Power BI, Tableau.
That seating chart is fine for a new hire. It is a bad architecture. pandas is still the interchange for sklearn and a pile of plotting code. Polars is the local engine when the frame is wide and the CPU has cores. DuckDB is SQL on files. Spark is a cluster. If you install all of them because a list numbered to 15, you own four runtimes and one messy notebook.
The FAQ in the same article is more useful than the ranking. What is the difference between pandas and Polars? Both are DataFrames; Polars is built for heavier local work. When DuckDB or Spark? DuckDB for fast analytical SQL on local data; Spark when the job leaves one machine. Which tools for LLM work? Hugging Face plus the rest of the deep-learning row. You knew that.
Python “still matters in 2026” is not news. It is the glue. Glue is not an engine. If your 2026 plan is “we use Python,” you have not chosen a frame library. You have chosen an interpreter.
Do not start a migration because Analytics Insight put Polars in the scale bucket. Start a migration because a job blew RAM or a scan read last year’s months for a September dashboard. That is the other URL.
If you need a default for a greenfield analysis repo: DuckDB or Polars for the heavy read, pandas at the edge where a library still wants it. Polars 2.0’s streaming default is the other local story. This week’s dated lab note is DuckDB on Parquet.
v1.5.x is partition pruning, not a mascot
The DuckDB Lab post is dated 2026-09-26. Version line: DuckDB v1.5.x. Test box: Linux, x86_64, 16 GB RAM. The example is the one every lake already claims and many notebooks ignore:
read_parquet('data/year=*/month=*/events.parquet')
WHERE year = 2026 AND month = 9
Comment in the post: DuckDB will only read the 2026-09 partition file instead of scanning all 33 months. That is Hive-style layout doing actual work. If your files are events_2024_through_now.parquet as one object, there is nothing to prune. The engine cannot skip a month that is not a directory.
Predicate pushdown is the second trick. A WHERE timestamp >= '2026-01-01' can be pushed into the Parquet read layer so column chunks that cannot match are not fully materialized. The post contrasts a naive read_parquet plus filter in Python with the SQL form that lets DuckDB push the predicate. If you SELECT * into a pandas DataFrame first, you paid for the scan. .df() after a tight SELECT is the cheap conversion.
Official docs are linked from the post. Use those for syntax. The lab page is a tutorial with affiliate VPS clutter in the chrome. Steal the query, not the hosting pitch.
33 months is a specific image. If your lake is four years of daily partitions, the skip is larger. If your lake is one folder of randomly named files, you do not have this feature. You have a mess. Layout is the project. DuckDB is the reader.
16 GB RAM is also a spec. DuckDB will still stream better than pandas on that box. It will not save you from SELECT * on an unpartitioned 80 GB dump. Projection (only the columns you name) is the other half of the skip. The listicle will not say that. The query will.
pandas is the adapter, not the scan
Keep pandas where the ecosystem is sticky: sklearn’s fit, a seaborn call, a teammate’s notebook that will not be rewritten this quarter. Convert at the edge. Do not load the lake with pd.read_parquet on the whole prefix “because we always have.”
Arrow is the peace treaty. pandas 2/3 with Arrow dtypes, Polars, DuckDB, Spark — they can pass batches without a CSV round trip. We have written the pandas-Arrow path before. The week’s increment is not a new dtype. It is a partition layout plus a reader that respects it.
If your job is a 20 MB CSV, pandas is fine. The tools list is written for that person and for the person with a lake, as if they were the same person. They are not. The 20 MB person should stop reading lists. The lake person should stop opening pandas first.
A useful rule from the FAQ, restated without the brochure: DuckDB when the question is SQL and the data is files. Polars when the question is a frame API and the CPU is bored. Spark when you already have a cluster and a reason. pandas when a library demands it. Four sentences. Fifteen tools was the long version.
Do not “standardize on pandas 3” and call it a 2026 stack. pandas 3 is a better adapter. It is still an adapter.
Measure once. Time pd.read_parquet on the prefix versus DuckDB read_parquet with the month filter. If they are within a second, your data is small and the list does not matter. If DuckDB returns before pandas allocates, you have the lab note in your repo. Check that time into the PR so the next listicle cannot undo it.
Spark is for leaving the box
Sharma puts Spark in the scale bucket with Polars and DuckDB. That is a category error. Polars and DuckDB are local. Spark is distributed. You turn on Spark when the partition skip is not enough because the month itself does not fit, or because the org already runs Spark and will not let a laptop be the warehouse.
If your “big data” is 12 GB of Parquet and a 32 GB workstation, Spark is a timeout with extra YAML. If your “big data” is a team that already submits to YARN or a cloud Spark SQL warehouse, stay there and still partition by month. Pruning is not a DuckDB exclusive. It is a layout exclusive. DuckDB is just the engine that makes the layout obvious on a laptop.
MLflow in the same bucket is observability for models, not for scans. Power BI and Tableau are other people’s SQL. If those tools query the lake, the same year= month= directories determine whether a VP dashboard scans 33 months. The BI tool will not save you from a bad prefix.
Hugging Face and PyTorch in a data-science tools list are how every 2026 roundup becomes an AI roundup. Fine if you train. Noise if you are here for a month of events. Do not pip install transformers to count rows.
What to install this week
If you already have DuckDB in the project, upgrade toward 1.5.x before you rewrite jobs. Then fix paths. Directories named year=2026/month=09/ are the feature. Cute filenames are not.
Rewrite the hottest notebook query as SQL against read_parquet with an explicit column list and a month predicate. Leave the pandas plot at the end on the small result.
If the notebook is Polars-first and happy, do not convert it to DuckDB because a list put both in the same bullet. They overlap. Pick one local engine per repo unless you have a reason.
If there is no layout, that is the sprint. Dumping a new year as a single file is how you make v1.5.x look ordinary.
Do not add Tableau to a Python repo because the list said so. Do not add TensorFlow because the list said so. The dated work this week is a WHERE clause that skips files.
Write the skip in the README with the 2026-09-26 lab URL and the Analytics Insight URL as the anti-pattern: seating charts versus a partition. New hires should read the query first.
Log the bytes read if your object store will tell you. A dashboard that “feels fast” can still pull last winter. Bytes are the review. Latency lies when the cache is warm.
If security wrapping blocks directory listing, pruning still works when the query supplies year and month as literals. It fails when those values are computed after a full glob in Python. Do the filter in SQL. That is the whole advanced section of the lab note, minus the ads.
What not to do with a 15-tool poster
Do not title an internal wiki “Our 2026 stack” and paste Sharma’s three buckets. Title it with the engine that reads production Parquet and the adapter you still need for sklearn.
Do not claim pandas is dead. The list still starts there because the ecosystem still starts there. Dead is a blog genre. Adapter is a job.
Do not claim DuckDB replaced Spark. The FAQ in the same listicle already split them. Believe the FAQ.
Do not copy the lab’s VPS aside into a production doc. The useful part is year and month and pushdown.
If your lake is Iceberg or DuckLake, you already have a catalog story from last week’s piece. Hive directories are the older cousin. Both beat a folder of mystery files. Mystery files are what make every engine look slow and every list look wise.
Run the September filter tomorrow on real files. If the process reads one month, you are done with the poster. If it reads 33, you are not having a pandas-versus-Polars debate. You are missing directories. Add the directories. Then argue about APIs if you still have time.
Discussion
Leave a comment
No comments yet
Be the first to start the conversation.