SlideShare uma empresa Scribd logo
1 de 11
Baixar para ler offline
Strategic  Advisory
Big  Data  – Cloud   -­‐ Analytics
Info
Strategy
Fishing  in  the  
big  data  lake
DATA  EXPLORATION  AND  DISCOVERY  ANALYTICS  
FOR  DEEPER  BUSINESS  INSIGHTS
InfoStrategy
What  is  a  “data  lake”
data  lake (plural data  lakes)
A  massive,  easily  accessible  data  repository  
built  on  (relatively)  inexpensive  computer  
hardware  for  storing  "big  data".  Unlike  data  marts,  
which  are  optimized  for  data  analysis  by  storing  only  some  
attributes  and  dropping  data  below  the  level  aggregation,  a  
data  lake  is  designed  to  retain  all  attributes,  
especially  so  when  you  do  not  yet  know  what  the  
scope  of  data  or  its  use  will  be.
http://en.wiktionary.org/wiki/data_lake
…  Enterprise  Data  Hub  sounds  too  boring   !
InfoStrategy
Optimise  business  through  insights
Insight
Action
Optimise
Move  a  metric
Change  a  product
Change  behaviour/process
Hindsight
Realtime
Foresight
Trusted  information
Act  on  insights  gained
Execute  theories
Measure
Outcomes
Sentiment
Feedback
Explore  datasets,  discover  correlations,  patterns.
Undiscovered  facts
Information  Value
Data  Volumes
Forecasting,  planning  &  trending
Statistical  Analysis
Operational  reporting,  SCADA  control
Alerts  &  Events
Historical  reporting, Proof  of  operation
Regulatory,  statutory,  financial
Uncover  previously  
unknown  facts  
from  enriched  data  
in  the  data  lake
InfoStrategy
Future  state  of  analytics
Strategic  Intent
To  improve  BI  and  Analytical  capabilities  to  a  level  where  organisations  are  able  to  
access  and  analyse  information  in  a  secure,  timely  and  cost-­‐effective  manner.
Gain  key  insights  to  optimise  the  operations  of  your  business,  predict  the  best  
possible  outcomes  for  growth,  new  opportunities,   and  competitive  advantage  
across  all  business  lines.
Mission  Statement
“Providing  advanced  analytics  capability  across  all  business  units,  empowering  our  
people  with  the    processes  and  supporting  technologies  to  exploit  our  information  
assets  for  business  benefit.”
Target  Operating  Model  will  deliver:
Rapid  access  to  data  to  uncover  new  facts  via  advanced  data  exploration  and  
discovery  analytics.
Clarity  of  who  is  responsible  and  accountable  for  maintaining  critical  information  
assets  via  a  well  structured  governance  and  engagement  model.
A  trusted  and  highly  secure  source  of  data  for  all  analytical  information  requirements  
via  a  data  quality  assurance  program.
Trawling  for  value  in  the  big  data  lake
InfoStrategy
‘Fish  stocks’  are  replenished  from  existing  and  future  
operational  systems  plus  external  sources
Core  
Transactional  Data  
“operational”
Management  
Reporting
Unstructured  &  
External  Data
“contextual”
Enterprise  Dashboards
Reporting
Consolidation
Data  ScientistsBusiness  AnalystsBusiness  UsersCustomers
Data  Extraction
Discovery  Analytics  
Platform
Visualisation
Analysis
Data  Preparation
Data  Collection
Operational  
Reporting
Operational  Dashboards
Real-­‐time  Reports
Alerts  &  Exceptions
Embedded  BI
Production   Data  Repository
“Data  Lake”
Information  Governance
Data  Management
Supplier  &  
Industry  Data
“comparative”
InfoStrategy
Consolidated
Management
Reporting
Operational
Supporting
Capability
Discovery
Analytics
To  meet  the  demand  for  rapid  access  to  information  
users  must  adopt  a  flexible  multi-­‐platform   architecture  
What  reporting  does  for  established  operations  …  discovery  analytics  does  for  new  business  development.
The  trend  within  industry  is  to  move  away  from  the  single-­‐platform  monolithic  data  warehouses  towards  a  physically  distributed  environment  
for  information  delivery.  Many  businesses  are  extending  their  data  warehouse  environments  to  include  new  standalone  data  platforms  that  
are  conducive  to  discovery  analytics.  A  holistic  view  is  maintained  via  a  common,  single  replicated  dataset  and  an  enterprise information  
management  program,  governing  delivery  and  access  to  key  information  (data  lake).
Source   Applications
ERP
CRM
HR
Finance
Telemetry
Geospatial  GIS
Documents
Email
Files
Real-­time  Data  
Capture
Cleansing
Loading
Data  Warehouse
Modelling
Relational  DW
Data  Marts
Analysis  Cubes
Analytics Delivery
Cloud-­based    Service  Model
Actuarial  
Applications
Event-­Based  
Applications
Reporting
Production  
Reporting
OLAP  Analytics
Ad  Hoc  Query
External
Data
Exploration  &  
Discovery
Metadata  Integration
Event  Processing Results
Detailed  Datasets Results  
Collection  and  blending Insights
Portal
PDF
Desktop
Guided  
Visualisation
Mobile  BI
Active  
Dashboards
Data  Replication
Historical Data  Preparation
Storytelling
Information  Governance
Operational  Reporting  
Dimensional  
Modelling
ProductioniseInsights
InfoStrategy
Principles:  Easier  access  information   to  discover  new  
facts  about  the  business.
◦ Described  as  a  ‘sandpit’  environment,  providing  the  ability  to  explore  and  discover  new  
facts  about  the  business,  it’s  members  and  customers,  partners  and  competitive  
pressures.
◦ Also  used  for  testing  a  hypothesis  or  running  scenarios  across  the  data
◦ Getting  answers  to  ‘one-­‐off’  questions  which  are  not  addressed  through  the  normal  
published,  scheduled  operational  reporting  channels
◦ Data  is  replicated  from  all  operational  systems  into  a  single  landing  area,  ensuring  
traceability  and  reconciliation  to  all  consuming  applications,  such  as  the  data  warehouse,  
analytical  application,  and  other  business  applications.
◦ Clearly  defined  critical  business  entities/records  are  synchronised  (or  Mastered)  across  
all  applications  eliminating  duplication  and  confusion.  Data  quality  attributes  are  defined  
and  managed  for  each  critical  business  entity.
◦ A  fully  integrated  Member/Customer  view  is  established  across  both  analytical  and  
transactional  applications.
◦ Using  the  replicated  data  to  build  more  dynamic  analytical  data  structures  for  scheduled  
production  reporting  and  ah-­‐hoc  analysis
◦ Provide  users  with  the  tools  to  access    and  analyse data,  freely  explore  current  and  new  
datasets,  and  visualise patterns  and  discoveries  to  gain  deep  insights.
Providing  business  users  with  direct  
access  to  data  to  meet  immediate  
information  needs  where  the  
accuracy  of  the  data  is  not  the  
primary  objective.  
Having  a  single  source  of  truth  
across  all  business  applications  at  
detailed  level  from  which  all  
information  requests  are  satisfied.
Improved  environment  for  more  
cost  effective  and  faster  business  
intelligence  delivery.
Provide  business   users  with  the  ability  to  access  production  information  directly,  collect  it  as  needed,  and  
prepare  the  data  for  analysis.  Exploring  the  data  to  uncover  previously   unknown  facts  about  the  business,   and  
sharing  those  facts  visually  with  others.  Enrich  production  data  with  external  “context”  to  extend  insights.
Key  Principles Description
InfoStrategy
Benefits  of  Discovery  Analytics  versus  traditional   data  
warehousing
Classic  Data  Warehouse  Issues Discovery  Analytics Benefit
Lengthy  IT  Backlog  and  lack  of  resources  to  extend the  
EDW  to  support  new  business  requirements.
Data  can  be  explored  and  analysed  outside  of the  EDW  
environment  before  it  is  put  into  production  use.
High  costs  of  supporting increasing  data  volumes  and  
new  types  of  data.
Data  can  be  filtered  and  transformed  before  it  is  loaded  
into  the  EDW
Lack  of  flexibility  in  the  EDW  data  model  to  support  
constantly changing  business  requirements.
Data  discovery  support  dynamic  schema  on  read  
approach  which reduces  the  need  for  detailed  up-­‐front  
modelling.
Need  to  have  data  quality  and  governance  processes  in  
place  before  user  can  access  the  EDW  data.
The  investigative  nature  of data  discovery  has  lower  data  
quality  and  governance  requirements
Growing  use  of  personal  data  marts to  overcome  IT  
barriers  and  the  performance  overheads  of  ad  hoc  
processing
The  flexibility  and  performance  of  data  discovery  
encourages  shared  use  of  data  and  analytics.
Recent  proof  of  concept  for  Discovery  Analytics  in  the  cloud  (AWS),  has  provided  some  
considerable  cost  &  time  savings  in  infrastructure  and  hosting,  viz.:
$55  per  day  to  host  a  960GB  data  warehouse  
$32  per  day  to  host  a  Data  Integration  server  AND  a  BI  server.
2.5  weeks  to  setup  POC  environment  and  start  analysis  and  visualising  results.
InfoStrategy
Discovery  Analytics  Target  POC  Architecture
Structured  
Data
Unstructured  
Data
ERP
Telemetry
Web/External
Replication  of  corporate  data,  enriched  with  external  data  and  
content,  available  in  a  centrally  available  and  scalable  repository  
ready  for  exploration,  discovery  and  predictive  analysis  to  gain  
deep  insights  and  actionable  results.
InfoStrategy
Fishing  safely  with  the  appropriate  life  vests  is  
important  too.
Security  and  data  management  standards  are  available
International  
Standard  on  
Assurance  
Engagements
Service  Organisation  
Control  framework
Federal  Information  
Management  
Security  Act
Payment  Card  
Industry  –Data  
Security  Standard
Federal  Information  
Processing  Standard
International  Standards  
Organisation  –
Information  Security  
Standard
Source:  Amazon  Web  Services
Info
Strategy
To  learn  more  about  how  InfoStrategy
can  help  you  develop  your  big  data  
strategy  to  solve  your  big  business  
problems,  or  to  arrange  a  Proof  of  
Concept,  please  contact  us  today  using  
the  details  below.
InfoStrategy Pty  Ltd
246  Oxford  St,  Balmoral
Queensland  4171
Australia
Tel:  +61  7  3151  2021
Email:  
contactus@infostrategy.com.au

Mais conteúdo relacionado

Mais procurados

Modernizing to a Cloud Data Architecture
Modernizing to a Cloud Data ArchitectureModernizing to a Cloud Data Architecture
Modernizing to a Cloud Data ArchitectureDatabricks
 
Is the traditional data warehouse dead?
Is the traditional data warehouse dead?Is the traditional data warehouse dead?
Is the traditional data warehouse dead?James Serra
 
Data Lakehouse, Data Mesh, and Data Fabric (r2)
Data Lakehouse, Data Mesh, and Data Fabric (r2)Data Lakehouse, Data Mesh, and Data Fabric (r2)
Data Lakehouse, Data Mesh, and Data Fabric (r2)James Serra
 
Big data architectures and the data lake
Big data architectures and the data lakeBig data architectures and the data lake
Big data architectures and the data lakeJames Serra
 
Power BI for Big Data and the New Look of Big Data Solutions
Power BI for Big Data and the New Look of Big Data SolutionsPower BI for Big Data and the New Look of Big Data Solutions
Power BI for Big Data and the New Look of Big Data SolutionsJames Serra
 
Data Lakehouse, Data Mesh, and Data Fabric (r1)
Data Lakehouse, Data Mesh, and Data Fabric (r1)Data Lakehouse, Data Mesh, and Data Fabric (r1)
Data Lakehouse, Data Mesh, and Data Fabric (r1)James Serra
 
Intro to Data Vault 2.0 on Snowflake
Intro to Data Vault 2.0 on SnowflakeIntro to Data Vault 2.0 on Snowflake
Intro to Data Vault 2.0 on SnowflakeKent Graziano
 
Data Architecture Best Practices for Advanced Analytics
Data Architecture Best Practices for Advanced AnalyticsData Architecture Best Practices for Advanced Analytics
Data Architecture Best Practices for Advanced AnalyticsDATAVERSITY
 
Data Lakehouse Symposium | Day 4
Data Lakehouse Symposium | Day 4Data Lakehouse Symposium | Day 4
Data Lakehouse Symposium | Day 4Databricks
 
DW Migration Webinar-March 2022.pptx
DW Migration Webinar-March 2022.pptxDW Migration Webinar-March 2022.pptx
DW Migration Webinar-March 2022.pptxDatabricks
 
Actionable Insights with AI - Snowflake for Data Science
Actionable Insights with AI - Snowflake for Data ScienceActionable Insights with AI - Snowflake for Data Science
Actionable Insights with AI - Snowflake for Data ScienceHarald Erb
 
Data Mesh for Dinner
Data Mesh for DinnerData Mesh for Dinner
Data Mesh for DinnerKent Graziano
 
Data Architecture Best Practices for Today’s Rapidly Changing Data Landscape
Data Architecture Best Practices for Today’s Rapidly Changing Data LandscapeData Architecture Best Practices for Today’s Rapidly Changing Data Landscape
Data Architecture Best Practices for Today’s Rapidly Changing Data LandscapeDATAVERSITY
 
Building a Data Strategy – Practical Steps for Aligning with Business Goals
Building a Data Strategy – Practical Steps for Aligning with Business GoalsBuilding a Data Strategy – Practical Steps for Aligning with Business Goals
Building a Data Strategy – Practical Steps for Aligning with Business GoalsDATAVERSITY
 
Microsoft Data Platform - What's included
Microsoft Data Platform - What's includedMicrosoft Data Platform - What's included
Microsoft Data Platform - What's includedJames Serra
 
Building Lakehouses on Delta Lake with SQL Analytics Primer
Building Lakehouses on Delta Lake with SQL Analytics PrimerBuilding Lakehouses on Delta Lake with SQL Analytics Primer
Building Lakehouses on Delta Lake with SQL Analytics PrimerDatabricks
 
Azure data platform overview
Azure data platform overviewAzure data platform overview
Azure data platform overviewJames Serra
 

Mais procurados (20)

From Data Warehouse to Lakehouse
From Data Warehouse to LakehouseFrom Data Warehouse to Lakehouse
From Data Warehouse to Lakehouse
 
Modernizing to a Cloud Data Architecture
Modernizing to a Cloud Data ArchitectureModernizing to a Cloud Data Architecture
Modernizing to a Cloud Data Architecture
 
Is the traditional data warehouse dead?
Is the traditional data warehouse dead?Is the traditional data warehouse dead?
Is the traditional data warehouse dead?
 
Data Lakehouse, Data Mesh, and Data Fabric (r2)
Data Lakehouse, Data Mesh, and Data Fabric (r2)Data Lakehouse, Data Mesh, and Data Fabric (r2)
Data Lakehouse, Data Mesh, and Data Fabric (r2)
 
Big data architectures and the data lake
Big data architectures and the data lakeBig data architectures and the data lake
Big data architectures and the data lake
 
Power BI for Big Data and the New Look of Big Data Solutions
Power BI for Big Data and the New Look of Big Data SolutionsPower BI for Big Data and the New Look of Big Data Solutions
Power BI for Big Data and the New Look of Big Data Solutions
 
Data Mesh
Data MeshData Mesh
Data Mesh
 
Data Lakehouse, Data Mesh, and Data Fabric (r1)
Data Lakehouse, Data Mesh, and Data Fabric (r1)Data Lakehouse, Data Mesh, and Data Fabric (r1)
Data Lakehouse, Data Mesh, and Data Fabric (r1)
 
Intro to Data Vault 2.0 on Snowflake
Intro to Data Vault 2.0 on SnowflakeIntro to Data Vault 2.0 on Snowflake
Intro to Data Vault 2.0 on Snowflake
 
Data Architecture Best Practices for Advanced Analytics
Data Architecture Best Practices for Advanced AnalyticsData Architecture Best Practices for Advanced Analytics
Data Architecture Best Practices for Advanced Analytics
 
Data Lakehouse Symposium | Day 4
Data Lakehouse Symposium | Day 4Data Lakehouse Symposium | Day 4
Data Lakehouse Symposium | Day 4
 
DW Migration Webinar-March 2022.pptx
DW Migration Webinar-March 2022.pptxDW Migration Webinar-March 2022.pptx
DW Migration Webinar-March 2022.pptx
 
Actionable Insights with AI - Snowflake for Data Science
Actionable Insights with AI - Snowflake for Data ScienceActionable Insights with AI - Snowflake for Data Science
Actionable Insights with AI - Snowflake for Data Science
 
Data Mesh for Dinner
Data Mesh for DinnerData Mesh for Dinner
Data Mesh for Dinner
 
Snowflake Overview
Snowflake OverviewSnowflake Overview
Snowflake Overview
 
Data Architecture Best Practices for Today’s Rapidly Changing Data Landscape
Data Architecture Best Practices for Today’s Rapidly Changing Data LandscapeData Architecture Best Practices for Today’s Rapidly Changing Data Landscape
Data Architecture Best Practices for Today’s Rapidly Changing Data Landscape
 
Building a Data Strategy – Practical Steps for Aligning with Business Goals
Building a Data Strategy – Practical Steps for Aligning with Business GoalsBuilding a Data Strategy – Practical Steps for Aligning with Business Goals
Building a Data Strategy – Practical Steps for Aligning with Business Goals
 
Microsoft Data Platform - What's included
Microsoft Data Platform - What's includedMicrosoft Data Platform - What's included
Microsoft Data Platform - What's included
 
Building Lakehouses on Delta Lake with SQL Analytics Primer
Building Lakehouses on Delta Lake with SQL Analytics PrimerBuilding Lakehouses on Delta Lake with SQL Analytics Primer
Building Lakehouses on Delta Lake with SQL Analytics Primer
 
Azure data platform overview
Azure data platform overviewAzure data platform overview
Azure data platform overview
 

Semelhante a Data lake benefits

intelligent-data-lake_executive-brief
intelligent-data-lake_executive-briefintelligent-data-lake_executive-brief
intelligent-data-lake_executive-briefLindy-Anne Botha
 
What Data Do You Have and Where is It?
What Data Do You Have and Where is It? What Data Do You Have and Where is It?
What Data Do You Have and Where is It? Caserta
 
BI Masterclass slides (Reference Architecture v3)
BI Masterclass slides (Reference Architecture v3)BI Masterclass slides (Reference Architecture v3)
BI Masterclass slides (Reference Architecture v3)Syaifuddin Ismail
 
Setting Up the Data Lake
Setting Up the Data LakeSetting Up the Data Lake
Setting Up the Data LakeCaserta
 
Datawarehousing
DatawarehousingDatawarehousing
Datawarehousingwork
 
Derfor skal du bruge en DataLake
Derfor skal du bruge en DataLakeDerfor skal du bruge en DataLake
Derfor skal du bruge en DataLakeMicrosoft
 
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...Why Your Data Science Architecture Should Include a Data Virtualization Tool ...
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...Denodo
 
BAR360 open data platform presentation at DAMA, Sydney
BAR360 open data platform presentation at DAMA, SydneyBAR360 open data platform presentation at DAMA, Sydney
BAR360 open data platform presentation at DAMA, SydneySai Paravastu
 
CS8091_BDA_Unit_I_Analytical_Architecture
CS8091_BDA_Unit_I_Analytical_ArchitectureCS8091_BDA_Unit_I_Analytical_Architecture
CS8091_BDA_Unit_I_Analytical_ArchitecturePalani Kumar
 
Big Data's Impact on the Enterprise
Big Data's Impact on the EnterpriseBig Data's Impact on the Enterprise
Big Data's Impact on the EnterpriseCaserta
 
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...Denodo
 
Dataware housing
Dataware housingDataware housing
Dataware housingwork
 
Data Virtualization. An Introduction (ASEAN)
Data Virtualization. An Introduction (ASEAN)Data Virtualization. An Introduction (ASEAN)
Data Virtualization. An Introduction (ASEAN)Denodo
 
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...DataScienceConferenc1
 
Overview of Business Intelligence
Overview of Business IntelligenceOverview of Business Intelligence
Overview of Business IntelligenceParthiv Dixit
 
Big Data Analytics and Machine Learning Document.docx
Big Data Analytics and Machine Learning Document.docxBig Data Analytics and Machine Learning Document.docx
Big Data Analytics and Machine Learning Document.docxZitin Technologies PVT LTD
 
What's New in Pentaho 7.0?
What's New in Pentaho 7.0?What's New in Pentaho 7.0?
What's New in Pentaho 7.0?Xpand IT
 
Big data journey to the cloud maz chaudhri 5.30.18
Big data journey to the cloud   maz chaudhri 5.30.18Big data journey to the cloud   maz chaudhri 5.30.18
Big data journey to the cloud maz chaudhri 5.30.18Cloudera, Inc.
 

Semelhante a Data lake benefits (20)

intelligent-data-lake_executive-brief
intelligent-data-lake_executive-briefintelligent-data-lake_executive-brief
intelligent-data-lake_executive-brief
 
What Data Do You Have and Where is It?
What Data Do You Have and Where is It? What Data Do You Have and Where is It?
What Data Do You Have and Where is It?
 
BI Masterclass slides (Reference Architecture v3)
BI Masterclass slides (Reference Architecture v3)BI Masterclass slides (Reference Architecture v3)
BI Masterclass slides (Reference Architecture v3)
 
Setting Up the Data Lake
Setting Up the Data LakeSetting Up the Data Lake
Setting Up the Data Lake
 
Datawarehousing
DatawarehousingDatawarehousing
Datawarehousing
 
Derfor skal du bruge en DataLake
Derfor skal du bruge en DataLakeDerfor skal du bruge en DataLake
Derfor skal du bruge en DataLake
 
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...Why Your Data Science Architecture Should Include a Data Virtualization Tool ...
Why Your Data Science Architecture Should Include a Data Virtualization Tool ...
 
Big data and oracle
Big data and oracleBig data and oracle
Big data and oracle
 
BAR360 open data platform presentation at DAMA, Sydney
BAR360 open data platform presentation at DAMA, SydneyBAR360 open data platform presentation at DAMA, Sydney
BAR360 open data platform presentation at DAMA, Sydney
 
CS8091_BDA_Unit_I_Analytical_Architecture
CS8091_BDA_Unit_I_Analytical_ArchitectureCS8091_BDA_Unit_I_Analytical_Architecture
CS8091_BDA_Unit_I_Analytical_Architecture
 
Big Data's Impact on the Enterprise
Big Data's Impact on the EnterpriseBig Data's Impact on the Enterprise
Big Data's Impact on the Enterprise
 
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...
How Data Virtualization Puts Enterprise Machine Learning Programs into Produc...
 
Dataware housing
Dataware housingDataware housing
Dataware housing
 
Data Virtualization. An Introduction (ASEAN)
Data Virtualization. An Introduction (ASEAN)Data Virtualization. An Introduction (ASEAN)
Data Virtualization. An Introduction (ASEAN)
 
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...
[DSC Europe 23] Milos Solujic - Data Lakehouse Revolutionizing Data Managemen...
 
Overview of Business Intelligence
Overview of Business IntelligenceOverview of Business Intelligence
Overview of Business Intelligence
 
Big Data Analytics and Machine Learning Document.docx
Big Data Analytics and Machine Learning Document.docxBig Data Analytics and Machine Learning Document.docx
Big Data Analytics and Machine Learning Document.docx
 
Machine Data Analytics
Machine Data AnalyticsMachine Data Analytics
Machine Data Analytics
 
What's New in Pentaho 7.0?
What's New in Pentaho 7.0?What's New in Pentaho 7.0?
What's New in Pentaho 7.0?
 
Big data journey to the cloud maz chaudhri 5.30.18
Big data journey to the cloud   maz chaudhri 5.30.18Big data journey to the cloud   maz chaudhri 5.30.18
Big data journey to the cloud maz chaudhri 5.30.18
 

Último

Call Girls in Defence Colony Delhi 💯Call Us 🔝8264348440🔝
Call Girls in Defence Colony Delhi 💯Call Us 🔝8264348440🔝Call Girls in Defence Colony Delhi 💯Call Us 🔝8264348440🔝
Call Girls in Defence Colony Delhi 💯Call Us 🔝8264348440🔝soniya singh
 
原版1:1定制南十字星大学毕业证(SCU毕业证)#文凭成绩单#真实留信学历认证永久存档
原版1:1定制南十字星大学毕业证(SCU毕业证)#文凭成绩单#真实留信学历认证永久存档原版1:1定制南十字星大学毕业证(SCU毕业证)#文凭成绩单#真实留信学历认证永久存档
原版1:1定制南十字星大学毕业证(SCU毕业证)#文凭成绩单#真实留信学历认证永久存档208367051
 
Top 5 Best Data Analytics Courses In Queens
Top 5 Best Data Analytics Courses In QueensTop 5 Best Data Analytics Courses In Queens
Top 5 Best Data Analytics Courses In Queensdataanalyticsqueen03
 
How we prevented account sharing with MFA
How we prevented account sharing with MFAHow we prevented account sharing with MFA
How we prevented account sharing with MFAAndrei Kaleshka
 
Multiple time frame trading analysis -brianshannon.pdf
Multiple time frame trading analysis -brianshannon.pdfMultiple time frame trading analysis -brianshannon.pdf
Multiple time frame trading analysis -brianshannon.pdfchwongval
 
INTERNSHIP ON PURBASHA COMPOSITE TEX LTD
INTERNSHIP ON PURBASHA COMPOSITE TEX LTDINTERNSHIP ON PURBASHA COMPOSITE TEX LTD
INTERNSHIP ON PURBASHA COMPOSITE TEX LTDRafezzaman
 
Consent & Privacy Signals on Google *Pixels* - MeasureCamp Amsterdam 2024
Consent & Privacy Signals on Google *Pixels* - MeasureCamp Amsterdam 2024Consent & Privacy Signals on Google *Pixels* - MeasureCamp Amsterdam 2024
Consent & Privacy Signals on Google *Pixels* - MeasureCamp Amsterdam 2024thyngster
 
Defining Constituents, Data Vizzes and Telling a Data Story
Defining Constituents, Data Vizzes and Telling a Data StoryDefining Constituents, Data Vizzes and Telling a Data Story
Defining Constituents, Data Vizzes and Telling a Data StoryJeremy Anderson
 
Call Girls In Dwarka 9654467111 Escorts Service
Call Girls In Dwarka 9654467111 Escorts ServiceCall Girls In Dwarka 9654467111 Escorts Service
Call Girls In Dwarka 9654467111 Escorts ServiceSapana Sha
 
Advanced Machine Learning for Business Professionals
Advanced Machine Learning for Business ProfessionalsAdvanced Machine Learning for Business Professionals
Advanced Machine Learning for Business ProfessionalsVICTOR MAESTRE RAMIREZ
 
Building on a FAIRly Strong Foundation to Connect Academic Research to Transl...
Building on a FAIRly Strong Foundation to Connect Academic Research to Transl...Building on a FAIRly Strong Foundation to Connect Academic Research to Transl...
Building on a FAIRly Strong Foundation to Connect Academic Research to Transl...Jack DiGiovanna
 
毕业文凭制作#回国入职#diploma#degree澳洲中央昆士兰大学毕业证成绩单pdf电子版制作修改#毕业文凭制作#回国入职#diploma#degree
毕业文凭制作#回国入职#diploma#degree澳洲中央昆士兰大学毕业证成绩单pdf电子版制作修改#毕业文凭制作#回国入职#diploma#degree毕业文凭制作#回国入职#diploma#degree澳洲中央昆士兰大学毕业证成绩单pdf电子版制作修改#毕业文凭制作#回国入职#diploma#degree
毕业文凭制作#回国入职#diploma#degree澳洲中央昆士兰大学毕业证成绩单pdf电子版制作修改#毕业文凭制作#回国入职#diploma#degreeyuu sss
 
Indian Call Girls in Abu Dhabi O5286O24O8 Call Girls in Abu Dhabi By Independ...
Indian Call Girls in Abu Dhabi O5286O24O8 Call Girls in Abu Dhabi By Independ...Indian Call Girls in Abu Dhabi O5286O24O8 Call Girls in Abu Dhabi By Independ...
Indian Call Girls in Abu Dhabi O5286O24O8 Call Girls in Abu Dhabi By Independ...dajasot375
 
dokumen.tips_chapter-4-transient-heat-conduction-mehmet-kanoglu.ppt
dokumen.tips_chapter-4-transient-heat-conduction-mehmet-kanoglu.pptdokumen.tips_chapter-4-transient-heat-conduction-mehmet-kanoglu.ppt
dokumen.tips_chapter-4-transient-heat-conduction-mehmet-kanoglu.pptSonatrach
 
PKS-TGC-1084-630 - Stage 1 Proposal.pptx
PKS-TGC-1084-630 - Stage 1 Proposal.pptxPKS-TGC-1084-630 - Stage 1 Proposal.pptx
PKS-TGC-1084-630 - Stage 1 Proposal.pptxPramod Kumar Srivastava
 
20240419 - Measurecamp Amsterdam - SAM.pdf
20240419 - Measurecamp Amsterdam - SAM.pdf20240419 - Measurecamp Amsterdam - SAM.pdf
20240419 - Measurecamp Amsterdam - SAM.pdfHuman37
 
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...Boston Institute of Analytics
 
Heart Disease Classification Report: A Data Analysis Project
Heart Disease Classification Report: A Data Analysis ProjectHeart Disease Classification Report: A Data Analysis Project
Heart Disease Classification Report: A Data Analysis ProjectBoston Institute of Analytics
 
办理学位证中佛罗里达大学毕业证,UCF成绩单原版一比一
办理学位证中佛罗里达大学毕业证,UCF成绩单原版一比一办理学位证中佛罗里达大学毕业证,UCF成绩单原版一比一
办理学位证中佛罗里达大学毕业证,UCF成绩单原版一比一F sss
 

Último (20)

Call Girls in Defence Colony Delhi 💯Call Us 🔝8264348440🔝
Call Girls in Defence Colony Delhi 💯Call Us 🔝8264348440🔝Call Girls in Defence Colony Delhi 💯Call Us 🔝8264348440🔝
Call Girls in Defence Colony Delhi 💯Call Us 🔝8264348440🔝
 
原版1:1定制南十字星大学毕业证(SCU毕业证)#文凭成绩单#真实留信学历认证永久存档
原版1:1定制南十字星大学毕业证(SCU毕业证)#文凭成绩单#真实留信学历认证永久存档原版1:1定制南十字星大学毕业证(SCU毕业证)#文凭成绩单#真实留信学历认证永久存档
原版1:1定制南十字星大学毕业证(SCU毕业证)#文凭成绩单#真实留信学历认证永久存档
 
Top 5 Best Data Analytics Courses In Queens
Top 5 Best Data Analytics Courses In QueensTop 5 Best Data Analytics Courses In Queens
Top 5 Best Data Analytics Courses In Queens
 
How we prevented account sharing with MFA
How we prevented account sharing with MFAHow we prevented account sharing with MFA
How we prevented account sharing with MFA
 
Multiple time frame trading analysis -brianshannon.pdf
Multiple time frame trading analysis -brianshannon.pdfMultiple time frame trading analysis -brianshannon.pdf
Multiple time frame trading analysis -brianshannon.pdf
 
INTERNSHIP ON PURBASHA COMPOSITE TEX LTD
INTERNSHIP ON PURBASHA COMPOSITE TEX LTDINTERNSHIP ON PURBASHA COMPOSITE TEX LTD
INTERNSHIP ON PURBASHA COMPOSITE TEX LTD
 
Consent & Privacy Signals on Google *Pixels* - MeasureCamp Amsterdam 2024
Consent & Privacy Signals on Google *Pixels* - MeasureCamp Amsterdam 2024Consent & Privacy Signals on Google *Pixels* - MeasureCamp Amsterdam 2024
Consent & Privacy Signals on Google *Pixels* - MeasureCamp Amsterdam 2024
 
Defining Constituents, Data Vizzes and Telling a Data Story
Defining Constituents, Data Vizzes and Telling a Data StoryDefining Constituents, Data Vizzes and Telling a Data Story
Defining Constituents, Data Vizzes and Telling a Data Story
 
Call Girls In Dwarka 9654467111 Escorts Service
Call Girls In Dwarka 9654467111 Escorts ServiceCall Girls In Dwarka 9654467111 Escorts Service
Call Girls In Dwarka 9654467111 Escorts Service
 
Advanced Machine Learning for Business Professionals
Advanced Machine Learning for Business ProfessionalsAdvanced Machine Learning for Business Professionals
Advanced Machine Learning for Business Professionals
 
Building on a FAIRly Strong Foundation to Connect Academic Research to Transl...
Building on a FAIRly Strong Foundation to Connect Academic Research to Transl...Building on a FAIRly Strong Foundation to Connect Academic Research to Transl...
Building on a FAIRly Strong Foundation to Connect Academic Research to Transl...
 
毕业文凭制作#回国入职#diploma#degree澳洲中央昆士兰大学毕业证成绩单pdf电子版制作修改#毕业文凭制作#回国入职#diploma#degree
毕业文凭制作#回国入职#diploma#degree澳洲中央昆士兰大学毕业证成绩单pdf电子版制作修改#毕业文凭制作#回国入职#diploma#degree毕业文凭制作#回国入职#diploma#degree澳洲中央昆士兰大学毕业证成绩单pdf电子版制作修改#毕业文凭制作#回国入职#diploma#degree
毕业文凭制作#回国入职#diploma#degree澳洲中央昆士兰大学毕业证成绩单pdf电子版制作修改#毕业文凭制作#回国入职#diploma#degree
 
Indian Call Girls in Abu Dhabi O5286O24O8 Call Girls in Abu Dhabi By Independ...
Indian Call Girls in Abu Dhabi O5286O24O8 Call Girls in Abu Dhabi By Independ...Indian Call Girls in Abu Dhabi O5286O24O8 Call Girls in Abu Dhabi By Independ...
Indian Call Girls in Abu Dhabi O5286O24O8 Call Girls in Abu Dhabi By Independ...
 
dokumen.tips_chapter-4-transient-heat-conduction-mehmet-kanoglu.ppt
dokumen.tips_chapter-4-transient-heat-conduction-mehmet-kanoglu.pptdokumen.tips_chapter-4-transient-heat-conduction-mehmet-kanoglu.ppt
dokumen.tips_chapter-4-transient-heat-conduction-mehmet-kanoglu.ppt
 
PKS-TGC-1084-630 - Stage 1 Proposal.pptx
PKS-TGC-1084-630 - Stage 1 Proposal.pptxPKS-TGC-1084-630 - Stage 1 Proposal.pptx
PKS-TGC-1084-630 - Stage 1 Proposal.pptx
 
20240419 - Measurecamp Amsterdam - SAM.pdf
20240419 - Measurecamp Amsterdam - SAM.pdf20240419 - Measurecamp Amsterdam - SAM.pdf
20240419 - Measurecamp Amsterdam - SAM.pdf
 
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...
NLP Data Science Project Presentation:Predicting Heart Disease with NLP Data ...
 
E-Commerce Order PredictionShraddha Kamble.pptx
E-Commerce Order PredictionShraddha Kamble.pptxE-Commerce Order PredictionShraddha Kamble.pptx
E-Commerce Order PredictionShraddha Kamble.pptx
 
Heart Disease Classification Report: A Data Analysis Project
Heart Disease Classification Report: A Data Analysis ProjectHeart Disease Classification Report: A Data Analysis Project
Heart Disease Classification Report: A Data Analysis Project
 
办理学位证中佛罗里达大学毕业证,UCF成绩单原版一比一
办理学位证中佛罗里达大学毕业证,UCF成绩单原版一比一办理学位证中佛罗里达大学毕业证,UCF成绩单原版一比一
办理学位证中佛罗里达大学毕业证,UCF成绩单原版一比一
 

Data lake benefits

  • 1. Strategic  Advisory Big  Data  – Cloud   -­‐ Analytics Info Strategy Fishing  in  the   big  data  lake DATA  EXPLORATION  AND  DISCOVERY  ANALYTICS   FOR  DEEPER  BUSINESS  INSIGHTS
  • 2. InfoStrategy What  is  a  “data  lake” data  lake (plural data  lakes) A  massive,  easily  accessible  data  repository   built  on  (relatively)  inexpensive  computer   hardware  for  storing  "big  data".  Unlike  data  marts,   which  are  optimized  for  data  analysis  by  storing  only  some   attributes  and  dropping  data  below  the  level  aggregation,  a   data  lake  is  designed  to  retain  all  attributes,   especially  so  when  you  do  not  yet  know  what  the   scope  of  data  or  its  use  will  be. http://en.wiktionary.org/wiki/data_lake …  Enterprise  Data  Hub  sounds  too  boring   !
  • 3. InfoStrategy Optimise  business  through  insights Insight Action Optimise Move  a  metric Change  a  product Change  behaviour/process Hindsight Realtime Foresight Trusted  information Act  on  insights  gained Execute  theories Measure Outcomes Sentiment Feedback Explore  datasets,  discover  correlations,  patterns. Undiscovered  facts Information  Value Data  Volumes Forecasting,  planning  &  trending Statistical  Analysis Operational  reporting,  SCADA  control Alerts  &  Events Historical  reporting, Proof  of  operation Regulatory,  statutory,  financial Uncover  previously   unknown  facts   from  enriched  data   in  the  data  lake
  • 4. InfoStrategy Future  state  of  analytics Strategic  Intent To  improve  BI  and  Analytical  capabilities  to  a  level  where  organisations  are  able  to   access  and  analyse  information  in  a  secure,  timely  and  cost-­‐effective  manner. Gain  key  insights  to  optimise  the  operations  of  your  business,  predict  the  best   possible  outcomes  for  growth,  new  opportunities,   and  competitive  advantage   across  all  business  lines. Mission  Statement “Providing  advanced  analytics  capability  across  all  business  units,  empowering  our   people  with  the    processes  and  supporting  technologies  to  exploit  our  information   assets  for  business  benefit.” Target  Operating  Model  will  deliver: Rapid  access  to  data  to  uncover  new  facts  via  advanced  data  exploration  and   discovery  analytics. Clarity  of  who  is  responsible  and  accountable  for  maintaining  critical  information   assets  via  a  well  structured  governance  and  engagement  model. A  trusted  and  highly  secure  source  of  data  for  all  analytical  information  requirements   via  a  data  quality  assurance  program. Trawling  for  value  in  the  big  data  lake
  • 5. InfoStrategy ‘Fish  stocks’  are  replenished  from  existing  and  future   operational  systems  plus  external  sources Core   Transactional  Data   “operational” Management   Reporting Unstructured  &   External  Data “contextual” Enterprise  Dashboards Reporting Consolidation Data  ScientistsBusiness  AnalystsBusiness  UsersCustomers Data  Extraction Discovery  Analytics   Platform Visualisation Analysis Data  Preparation Data  Collection Operational   Reporting Operational  Dashboards Real-­‐time  Reports Alerts  &  Exceptions Embedded  BI Production   Data  Repository “Data  Lake” Information  Governance Data  Management Supplier  &   Industry  Data “comparative”
  • 6. InfoStrategy Consolidated Management Reporting Operational Supporting Capability Discovery Analytics To  meet  the  demand  for  rapid  access  to  information   users  must  adopt  a  flexible  multi-­‐platform   architecture   What  reporting  does  for  established  operations  …  discovery  analytics  does  for  new  business  development. The  trend  within  industry  is  to  move  away  from  the  single-­‐platform  monolithic  data  warehouses  towards  a  physically  distributed  environment   for  information  delivery.  Many  businesses  are  extending  their  data  warehouse  environments  to  include  new  standalone  data  platforms  that   are  conducive  to  discovery  analytics.  A  holistic  view  is  maintained  via  a  common,  single  replicated  dataset  and  an  enterprise information   management  program,  governing  delivery  and  access  to  key  information  (data  lake). Source   Applications ERP CRM HR Finance Telemetry Geospatial  GIS Documents Email Files Real-­time  Data   Capture Cleansing Loading Data  Warehouse Modelling Relational  DW Data  Marts Analysis  Cubes Analytics Delivery Cloud-­based    Service  Model Actuarial   Applications Event-­Based   Applications Reporting Production   Reporting OLAP  Analytics Ad  Hoc  Query External Data Exploration  &   Discovery Metadata  Integration Event  Processing Results Detailed  Datasets Results   Collection  and  blending Insights Portal PDF Desktop Guided   Visualisation Mobile  BI Active   Dashboards Data  Replication Historical Data  Preparation Storytelling Information  Governance Operational  Reporting   Dimensional   Modelling ProductioniseInsights
  • 7. InfoStrategy Principles:  Easier  access  information   to  discover  new   facts  about  the  business. ◦ Described  as  a  ‘sandpit’  environment,  providing  the  ability  to  explore  and  discover  new   facts  about  the  business,  it’s  members  and  customers,  partners  and  competitive   pressures. ◦ Also  used  for  testing  a  hypothesis  or  running  scenarios  across  the  data ◦ Getting  answers  to  ‘one-­‐off’  questions  which  are  not  addressed  through  the  normal   published,  scheduled  operational  reporting  channels ◦ Data  is  replicated  from  all  operational  systems  into  a  single  landing  area,  ensuring   traceability  and  reconciliation  to  all  consuming  applications,  such  as  the  data  warehouse,   analytical  application,  and  other  business  applications. ◦ Clearly  defined  critical  business  entities/records  are  synchronised  (or  Mastered)  across   all  applications  eliminating  duplication  and  confusion.  Data  quality  attributes  are  defined   and  managed  for  each  critical  business  entity. ◦ A  fully  integrated  Member/Customer  view  is  established  across  both  analytical  and   transactional  applications. ◦ Using  the  replicated  data  to  build  more  dynamic  analytical  data  structures  for  scheduled   production  reporting  and  ah-­‐hoc  analysis ◦ Provide  users  with  the  tools  to  access    and  analyse data,  freely  explore  current  and  new   datasets,  and  visualise patterns  and  discoveries  to  gain  deep  insights. Providing  business  users  with  direct   access  to  data  to  meet  immediate   information  needs  where  the   accuracy  of  the  data  is  not  the   primary  objective.   Having  a  single  source  of  truth   across  all  business  applications  at   detailed  level  from  which  all   information  requests  are  satisfied. Improved  environment  for  more   cost  effective  and  faster  business   intelligence  delivery. Provide  business   users  with  the  ability  to  access  production  information  directly,  collect  it  as  needed,  and   prepare  the  data  for  analysis.  Exploring  the  data  to  uncover  previously   unknown  facts  about  the  business,   and   sharing  those  facts  visually  with  others.  Enrich  production  data  with  external  “context”  to  extend  insights. Key  Principles Description
  • 8. InfoStrategy Benefits  of  Discovery  Analytics  versus  traditional   data   warehousing Classic  Data  Warehouse  Issues Discovery  Analytics Benefit Lengthy  IT  Backlog  and  lack  of  resources  to  extend the   EDW  to  support  new  business  requirements. Data  can  be  explored  and  analysed  outside  of the  EDW   environment  before  it  is  put  into  production  use. High  costs  of  supporting increasing  data  volumes  and   new  types  of  data. Data  can  be  filtered  and  transformed  before  it  is  loaded   into  the  EDW Lack  of  flexibility  in  the  EDW  data  model  to  support   constantly changing  business  requirements. Data  discovery  support  dynamic  schema  on  read   approach  which reduces  the  need  for  detailed  up-­‐front   modelling. Need  to  have  data  quality  and  governance  processes  in   place  before  user  can  access  the  EDW  data. The  investigative  nature  of data  discovery  has  lower  data   quality  and  governance  requirements Growing  use  of  personal  data  marts to  overcome  IT   barriers  and  the  performance  overheads  of  ad  hoc   processing The  flexibility  and  performance  of  data  discovery   encourages  shared  use  of  data  and  analytics. Recent  proof  of  concept  for  Discovery  Analytics  in  the  cloud  (AWS),  has  provided  some   considerable  cost  &  time  savings  in  infrastructure  and  hosting,  viz.: $55  per  day  to  host  a  960GB  data  warehouse   $32  per  day  to  host  a  Data  Integration  server  AND  a  BI  server. 2.5  weeks  to  setup  POC  environment  and  start  analysis  and  visualising  results.
  • 9. InfoStrategy Discovery  Analytics  Target  POC  Architecture Structured   Data Unstructured   Data ERP Telemetry Web/External Replication  of  corporate  data,  enriched  with  external  data  and   content,  available  in  a  centrally  available  and  scalable  repository   ready  for  exploration,  discovery  and  predictive  analysis  to  gain   deep  insights  and  actionable  results.
  • 10. InfoStrategy Fishing  safely  with  the  appropriate  life  vests  is   important  too. Security  and  data  management  standards  are  available International   Standard  on   Assurance   Engagements Service  Organisation   Control  framework Federal  Information   Management   Security  Act Payment  Card   Industry  –Data   Security  Standard Federal  Information   Processing  Standard International  Standards   Organisation  – Information  Security   Standard Source:  Amazon  Web  Services
  • 11. Info Strategy To  learn  more  about  how  InfoStrategy can  help  you  develop  your  big  data   strategy  to  solve  your  big  business   problems,  or  to  arrange  a  Proof  of   Concept,  please  contact  us  today  using   the  details  below. InfoStrategy Pty  Ltd 246  Oxford  St,  Balmoral Queensland  4171 Australia Tel:  +61  7  3151  2021 Email:   contactus@infostrategy.com.au