
Deterministic infrastructure testing and probabilistic models are incongruent, yet must be reconciled. Model evaluations are the answer to this question, yet do not behave as unit tests or smoke tests because they require a large body of data that represents the phenomena being modeled. To make systems reliable with unwieldy, ever-changing data distributions, product-scope evaluations of machine learning systems, beyond the individual models, are a must.
Jack Sullivan is a machine learning engineer with GetReal Security in Austin, with a focus on enhancing identity verification for all. His experience spans startups, consulting, and large, federally-funded scientific organizations in machine learning and applied physics research, as well as platform and DevOps engineering.