Home - Confessions of a Data Guy

Data, Data Engineering, Python, Uncategorized

Big Data File Showdown – Avro vs Parquet with Python.

There comes a point in the life of every data person that we have to graduate from csv files. At a certain point the data becomes big enough or we hear talk on the street about other file formats. Apache Parquet and Apache Avro are two of those formats that been coming up more with […]

April 5, 2020

Data, Data Engineering, Machine Learning, Python

Challenges of Machine Learning Pipelines at Scale… When You Don’t Work at Google.

ml pipelines Building Machine Learning (ML) pipelines with big data is hard enough, and it doesn’t take much of a curve ball to make it a nightmare. Most of what you will read online are tutorials on how to take a few CSV files and run them through some sklearn package. If you are lucky, […]

March 14, 2020

Data, Data Engineering, Python, Uncategorized

Apache Airflow for Data Engineers

On again, off again. I feel like that is the best way to describe Apache Airflow. It started out around 2014 at Airbnb and has been steadily gaining traction and usage ever since, albeit slowly. I still believe that Airflow is very underutilized in the data engineering community as a whole, most everyone has heard […]

January 11, 2020

Data Engineering, Python, SQL

Introduction to Postgres with Python

If there was ever a match made in heaven, it’s using Python and Postgres together. They were made for each other. Both are fun and easy to use, addicting, both have so many surprises and hidden gems. Like Gandalf and Frodo, the two just go together. Today I want to go through the basics of […]

December 30, 2019

Data, Data Engineering, Python

Exploring ElasticSearch with Python

What’s Elasticsearch precious? I feel like Gollum when confronted by taters. Elasticsearch has been around for awhile now, based on Lucene, it’s become a well known name in the field of text and semi structured data storage, analysis and retrieve category. Even though it’s popular enough to get name recognition I’ve rarely run across it […]

December 17, 2019

Ramblings

Approaching Software as a Craft, then as a Engineer.

Craft first, engineering second. There’s probably a lot of software programmers, developers, and engineers who will take issue with this. That’s kinda the point. Software should be approached as a craft first, then a engineering problem second. There are so many ways this is true, it’s going to be hard to touch them all. I […]

December 9, 2019

Data, Data Engineering, Python

3 (Or More) Ways to Open a CSV in Python

Ah. What a classic. The one piece of code that I end up writing over and over again, you would think I would have stashed it away by now. Not going to lie I usually have to Google it, while thinking, is this the right way? Should I just open the csv file and iterate […]

November 27, 2019

Ramblings

How Smart Engineers Create Bad Software

You ever wonder how a room full of what appears to be smart engineers manage to build software that doesn’t work? Given more time and money, it appears to only get worse or no better. It doesn’t make that much sense does it? As someone who writes software it’s hard to see how bugs that […]

November 25, 2019

Data, Data Engineering, Geospatial, Python

Thunderdome for Geospatial Tools in Python. It’s to the Death.

It’s a fight to the death people… that’s why it’s called Thunderdome. This will be no different. Last time we talked about the very basics of the strange world of geo-spatial tools for data engineering. The next most obvious thing do of course is to see what tool is the best. By best I mean […]

November 23, 2019

Data, Geospatial, Python

Gentle Introduction to Geospatial for Data Engineers

What does a data engineer need to know about working with geospatial data? I’m going to give my two cents on what is and is not important. First, prepare to be annoyed as you will most likely spend hours debugging strange and not obvious errors and bugs. You should run screaming the other way, but […]

November 10, 2019

Big Data File Showdown – Avro vs Parquet with Python.

Challenges of Machine Learning Pipelines at Scale… When You Don’t Work at Google.

Apache Airflow for Data Engineers

Introduction to Postgres with Python

Exploring ElasticSearch with Python

Approaching Software as a Craft, then as a Engineer.

3 (Or More) Ways to Open a CSV in Python

How Smart Engineers Create Bad Software

Thunderdome for Geospatial Tools in Python. It’s to the Death.

Gentle Introduction to Geospatial for Data Engineers

Interesting links

Pages

Categories

Archive