
Some of the largest outages on the internet can be traced back not only to changes in code, but also how the code changed underlying data models. Through countless discussions with software engineers, many noted the importance of the underlying data model for quality development, yet also highlighted the lack of incentives (or outright discouragement) by leadership to put in the extra effort to maintain it. Even more troubling, not only are applications impacted by data, but also downstream consumers within the business are taking major dependencies on the output of this data for business-critical workflows-- unbeknownst to the upstream engineers producing the data (i.e., shadow dependencies). In this talk, we highlight this growing problem, why engineer leadership is paying more attention to the risk of data, and how to surface and prevent these issues within the CI/CD workflow via an emerging pattern called "data contracts."
Chad Sanderson is a data leader and CEO of Gable.ai, where he focuses on building platforms that improve collaboration and data quality at scale. With a background in journalism from Georgia Southern University, he brings a strong emphasis on storytelling, communication, and product thinking to complex data challenges. He also leads one of the fastest-growing data quality communities, Data Quality Camp, and is an author working on a guide to data contracts with O’Reilly.