Home - Confessions of a Data Guy

Using Rust to write a Data Pipeline. Thoughts. Musings.

Rust has been on my mind a lot lately, probably because of Data Engineering boredom, watching Spark clusters chug along like some medieval farm worker endlessly trudging through the muck and mire of life. Maybe Rust has breathed some life back into my stagnant soul, reminding me there is a big world out there, […]

January 13, 2023

Big Data, Data, Data Engineering, Data Warehousing

Simplify Delta Lake Complexity with mack.

Anyone who’s been roaming around the forest of Data Engineering has probably run into many of the newish tools that have been growing rapidly around the concepts of Data Warehouses, Data Lakes, and Lake Houses … the merging of the old relational database functionality with TB and PB level cloud-based file storage systems. Tools like […]

January 12, 2023

Data Engineering, Ramblings

I asked ChatGPT to write a blog post about Data Engineering. Here it is.

Data engineering is a vital field within the realm of data science that focuses on the practical aspects of collecting, storing, and processing large amounts of data. It involves designing and building the infrastructure to store and process data, as well as developing the tools and systems to extract valuable insights and knowledge from that […]

December 30, 2022

Big Data, Data, Data Engineering, Python

What is Apache Arrow? Asking for a friend.

We’ve all been in that spot, especially in tech. You wanted to fit in, be cool, and look smart, so you didn’t ask any questions. And now it’s too late. You’re stuck. Now you simply can’t ask … you’re too afraid. I get it. Apache Arrow is probably one of those things. It keeps popping […]

December 27, 2022

Big Data, Data, Data Engineering, Python, Ramblings, Rust

Dataframe Showdown – Polars vs Spark vs Pandas vs DataFusion. Guess who wins?

There once was a day when no one used DataFrames that much. Back before Spark had really gone mainstream, Data Scientists were still plinking around with Pandas a lot. My My, what would your mother say? How things have changed. Now everyone wants a piece of the DataFrame pie. I mean it tastes so good, […]

December 10, 2022

Big Data, Data, Data Engineering, Ramblings

Why Data Migrations Suck.

I’ve often wondered what purgatory would be like, doing penance for millennia into eternity. It would probably be doing data migrations. I suppose they are not all that dissimilar from normal software migrations, but there are a few things that make data migrations a little more horrible and soul-sucking. Data migrations are able to slow […]

December 5, 2022

Big Data, Data, Data Engineering, Data Warehousing, Ramblings

A Tale of Betrayal and Heartbreak – Databricks Workflows and Jobs.

Nothing captures the imagination and heart like a tale of betrayal and heartbreak, and that is a tale I want to bring to you today. It’s a tale of Databricks Workflows and Jobs, version changes, new features, API’s, and insidious little hidden gems that will make you pull your hair out when you find them. […]

December 1, 2022

Data, Data Engineering, Data Quality, Ramblings

A Diatribe against Data Contracts and their Abuses.

Ok, so I don’t really mean all that. Or do I? I have no idea what the future holds. Sometimes it’s easy to pick out the winners, like Databricks and Snowflake, you can see, feel, and taste the results of those data products, a delicious and delectable bounty to feast upon. Other things are harder […]

November 16, 2022

Big Data, Data, Data Engineering, Data Warehousing

Introduction to Historical Loads – for Data Engineers.

There are probably few things in life that will strike more fear and tumult in the heart of the Data Engineer than historical loads. You know, on the surface it seems like such an innocent thing. How could it possibly be, just take a bunch of data stored somewhere and shove it into a table. […]

November 5, 2022

Big Data, Data, Data Engineering, Data Warehousing, Rust

Delta Lake without Spark (delta-rs). Innovation, cost savings, and other such matters.

The intersection of Big Data and Not Big Data. An interesting topic of late that has been rattling around in my overcrowded head is the idea of Big Data vs Not Big Data, and the intersection thereof. I’ve been thinking about SAAS vendors, the Modern Data Stack, costs, and innovation. A great real-life example of […]

October 25, 2022

Using Rust to write a Data Pipeline. Thoughts. Musings.

Simplify Delta Lake Complexity with mack.

I asked ChatGPT to write a blog post about Data Engineering. Here it is.

What is Apache Arrow? Asking for a friend.

Dataframe Showdown – Polars vs Spark vs Pandas vs DataFusion. Guess who wins?

Why Data Migrations Suck.

A Tale of Betrayal and Heartbreak – Databricks Workflows and Jobs.

A Diatribe against Data Contracts and their Abuses.

Introduction to Historical Loads – for Data Engineers.

Delta Lake without Spark (delta-rs). Innovation, cost savings, and other such matters.

Interesting links

Pages

Categories

Archive