Provenance in Production-Grade Machine Learning

@AnandSampat +
Provenance in Production-Grade
Machine Learning
Talk
Santa Clara Convention Center

@AnandSampat +
Anand Sampat
CEO & Co-founder, Datmo
@AnandSampat

@AnandSampat +
Talk Outline
1. Rise of AI / ML in the Enterprise
2. Unique challenges of AI
3. Provenance, Reliability, and Efficiency
4. How Datmo bridges the gap

@AnandSampat +
Demand for Talent is Increasing
Today
Data Scientists: 48k
https://www.pwc.com/us/en/library/
data-science-and-analytics.html
Tomorrow
Data Engineers: 558k
http://www.mckinsey.com/business-functions/mckinsey-analytics/our-
insights/the-age-of-analytics-competing-in-a-data-driven-world

@AnandSampat +
Supply is Limited, but it’s growing
https://github.com

@AnandSampat +
QoD’s == Quantitative Oriented Developers
Artiﬁcial IntelligenceData Science Machine Learning

@TheNickWalsh +
Big Header
Section Header
Here are a bunch of words that will be
used to describe something. I’m
typing a bunch of words to fill up the
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +@TheNickWalsh
Am I a QoD?

@AnandSampat +
https://blog.datmo.io/demystifying-the-ml-ai-and-data-science-development-
ecosystem-part-1-build-76c6d4911d07

@AnandSampat +
https://blog.datmo.io/demystifying-the-ml-ai-and-data-science-development-
ecosystem-part-1-build-76c6d4911d07
+ Deployment! 
+ Post-Deployment!
(DevOps!)

@AnandSampat +
It’s time to talk about MLOps
https://papers.nips.cc/paper/5656-hidden-technical-debt-in-machine-learning-
systems.pdf

@AnandSampat +
MLOps: The Elephant in the Room
https://papers.nips.cc/paper/5656-hidden-technical-debt-in-machine-learning-
systems.pdf

@AnandSampat +
ML systems have a special capacity for incurring
technical debt, because they have all of the
maintenance problems of traditional code plus an
additional set of ML-specific issues. This debt may be
difficult to detect because it exists at the system level.
“
— Google (Sculley et. al, 2015)

@AnandSampat +
Typical methods for paying down code level
technical debt are not sufficient to address
ML-specific technical debt at the system level.
“
— Google (Sculley et. al, 2015)

@AnandSampat +
http://eng.uber.com/wp-content/uploads/2017/09/image8.png
Here’s where traditional tools fall short

@AnandSampat +
https://eng.uber.com/michelangelo/
https://code.facebook.com/posts/1072626246134461/
introducing-fblearner-flow-facebook-s-ai-backbone/

@AnandSampat +
As for everyone else?

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Provenance:
Model and Workflow
Reproducibility

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Problem: Model
reproduction is tough
- Configurations & Metrics
- Traditional SCM tools (like Git) do a
good job of tracking changes
between code snippets but
overlook machine learning
parameters and scoring metrics
- Dependencies
- Hardware Configuration
- GPU Setup/CUDA
- OS-level settings/programs
- How can you install packages
without a package manager?

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Solution: Tracking and
Containerization
- Track your configurations
and metrics in 1 place
- With containers, you can
write build files that enable
you to enumerate
everything required to
reproduce a given system
state
Problem: Model
reproduction is tough
- Configurations & Metrics
- Traditional SCM tools (like Git) do a
good job of tracking changes
between code snippets but
overlook machine learning
parameters and scoring metrics
- Dependencies
- Hardware Configuration
- GPU Setup/CUDA
- OS-level settings/programs
- How can you install packages
without a package manager?

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Example 1: “Offline” Logging (bad)

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Example 2: Online Logging with Visualizable
Metrics (good)
Unfortunately, TensorBoard
is only available
for TensorFlow!

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Example 3: Docker and Dockerfiles

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Reliability:
Peace of Mind

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
- Traditional software build tools
overlook model scoring and
metrics and thus do not check
builds for these metrics
- Traditional software
deployment don’t take into
account the nuances of
machine learning models
Problem: Builds and
deployments don’t account
for machine learning

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Solution:
Builds and Deployment
with machine learning
metrics
- Set scoring thresholds for
validation metrics of
models for builds
- Deploy your machine
learning as micro services
which can be updated on
a different schedule from
the main application.
Problem: Builds and
deployments don’t account
for machine learning
- Traditional software build tools
overlook model scoring and
metrics and thus do not check
builds for these metrics
- Traditional software
deployment don’t take into
account the nuances of
machine learning models

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Efficiency:
Reduce the time to success

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Problem: Disjoint tools
slow down iteration
- Software tools are not built to
iterate on machine learning
algorithms
- Machine learning does not
follow the same build schedule
as your main application

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Solution: A/B testing,
continuous deployment, and
automation
- A/B testing models enables
quick performance
comparisons to identify the
best parameters
- Continuous deployment
ensures that deployed
models work as expected
- Automation enables
triggers to create actions
Problem: Disjoint tools
slow down iteration
- Software tools are not built to
iterate on machine learning
algorithms
- Machine learning does not
follow the same build schedule
as your main application

@AnandSampat +
What is Datmo?
Datmo is a unified platform for ML, AI, and Data
Science developers. Datmo’s free Community
Edition enables model version control, easy
environment handling, and reproducing results
through the power of snapshots. Datmo
Enterprise leverages snapshots to enable
reliable builds, quick deployments, efficient A/
B testing and continuous delivery of analytics
workflows and models

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
Provenance: Datmo CE

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
- Snapshots - Model versions which combine code, files,
environments, configurations, and performance metrics
- Runnable Anywhere - The tool can be run on any server to
enable you to move your models freely between servers and
share them with colleagues
Datmo CE

@AnandSampat +
What are Datmo Snapshots?
Code
Environment
Configuration
Files*
Metrics

@AnandSampat +
Why are they important?
Environment
Configuration
Metrics
Datmo Snapshots
Git Commits
Code
Files*

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
GUI to View Snapshots

@AnandSampat +
How will it help?
Datmo leverages containers to quickly
spin up perfectly reproducible
developer environments. It tracks this
environment, along with model
metadata inside of snapshots.

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
Reliability: Datmo EE

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
- Builds - Model versions with Snapshot can be built by adding
validation tests that track your performance metrics
- Deployment - can be pushed as microservices so you can
update them on a different schedule from the rest of your main
application
Datmo EE
(Builds and Deployment)

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
@AnandSampat +
Deployment:
Containerization

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
Efficiency: Datmo EE

@TheNickWalsh +
Big Header
Section Header
box.
Medium header with
A lot of words
Caption
Subtitle
- A/B Testing — enables you to deploy a few microservices in
parallel which let’s you compare algorithms
- Continuous Deployment — enables you to update your builds
with tests that ensure your validation metrics meet your threshold
- Automation — Create triggers and actions to retrain your models
with new data, update your models frequently, or ensure you are
always in the know when models aren’t working.
Datmo EE
(A/B Testing, Continuous Deployment, Automation)

@AnandSampat +
Datmo CE + EE
Make ML Ops and workflows
manageable and simple, not
completely abstracted away.
Reduce the amount of glue code
so that people can have more
robust pipelines.

@AnandSampat +
1. AI applications are growing day-by-day. These
technologies require new capabilities
Key Takeaways
2. Provenance, Reliability, and Efficiency are required
for any production system — ML is no different
3. Datmo CE and EE provide full provenance, reliability,
and efficiency through snapshots which enable builds,
deployments, A/B testing and continuous delivery

@AnandSampat +
2015 NIPS Paper from Google
https://papers.nips.cc/paper/5656-hidden-
technical-debt-in-machine-learning-systems.pdf

@AnandSampat +
Learn More about Us at our Blog
https://blog.datmo.com/

@AnandSampat +
Check out our Product Pages
https://datmo.com/enterprisehttps://datmo.com/community

@AnandSampat +
Full Slides Available at:
http://bit.ly/global-ai-conf-provenance

Provenance in Production-Grade Machine Learning

Recomendados

Recomendados

Mais conteúdo relacionado

Mais procurados

Mais procurados (20)

Semelhante a Provenance in Production-Grade Machine Learning

Semelhante a Provenance in Production-Grade Machine Learning (20)

Último

Último (20)

Provenance in Production-Grade Machine Learning