Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Each session lasts 2 hours
Day-1: Session -1: Business Perspective on Big Data BI in Government
- Case studies from NIH, DoE
- Adoption rates of Big Data in government agencies and strategies for aligning future operations with Big Data Predictive Analytics
- Broad application areas in DoD, NSA, IRS, USDA, etc.
- Integration of Big Data with legacy data systems
- Fundamental understanding of enabling technologies in predictive analytics
- Data Integration & Dashboard visualization
- Fraud management
- Generation of Business Rules & Fraud detection
- Threat detection and profiling
- Cost-benefit analysis for Big Data implementation
Day-1: Session-2 : Introduction to Big Data-1
- Core characteristics of Big Data: volume, variety, velocity, and veracity. MPP architecture for handling volume.
- Data Warehouses – static schema, slowly evolving datasets
- MPP Databases such as Greenplum, Exadata, Teradata, Netezza, Vertica, etc.
- Hadoop Based Solutions – no restrictions on dataset structure.
- Typical pattern: HDFS, MapReduce (crunch), retrieval from HDFS
- Batch processing – suited for analytical/non-interactive tasks
- Volume handling: CEP streaming data
- Common choices – CEP products (e.g. Infostreams, Apama, MarkLogic, etc)
- Less production-ready options – Storm/S4
- NoSQL Databases – (columnar and key-value): Best suited as an analytical adjunct to data warehouses/databases
Day-1 : Session -3 : Introduction to Big Data-2
NoSQL Solutions
- KV Store - Keyspace, Flare, SchemaFree, RAMCloud, Oracle NoSQL Database (OnDB)
- KV Store - Dynamo, Voldemort, Dynomite, SubRecord, Mo8onDb, DovetailDB
- KV Store (Hierarchical) - GT.m, Cache
- KV Store (Ordered) - TokyoTyrant, Lightcloud, NMDB, Luxio, MemcacheDB, Actord
- KV Cache - Memcached, Repcached, Coherence, Infinispan, EXtremeScale, JBossCache, Velocity, Terracoqua
- Tuple Store - Gigaspaces, Coord, Apache River
- Object Database - ZopeDB, DB40, Shoal
- Document Store - CouchDB, Cloudant, Couchbase, MongoDB, Jackrabbit, XML-Databases, ThruDB, CloudKit, Prsevere, Riak-Basho, Scalaris
- Wide Columnar Store - BigTable, HBase, Apache Cassandra, Hypertable, KAI, OpenNeptune, Qbase, KDI
Data Varieties: Introduction to Data Cleaning Challenges in Big Data
- RDBMS – static structure/schema, does not foster an agile, exploratory environment.
- NoSQL – semi-structured, sufficient structure to store data without a predefined exact schema
- Data cleaning issues
Day-1 : Session-4 : Big Data Introduction-3 : Hadoop
- Criteria for selecting Hadoop
- STRUCTURED - Enterprise data warehouses/databases can store massive data (at a cost) but impose structure (limiting active exploration)
- SEMI-STRUCTURED data – challenging to handle with traditional solutions (DW/DB)
- Warehousing data = significant effort and static even after implementation
- For variety & volume of data, processed on commodity hardware – HADOOP
- Commodity Hardware needed to establish a Hadoop Cluster
Introduction to Map Reduce /HDFS
- MapReduce – distributing computing tasks across multiple servers
- HDFS – making data locally available for computing processes (with redundancy)
- Data – can be unstructured/schema-less (unlike RDBMS)
- Developer responsibility to interpret data
- Programming MapReduce = working with Java (pros/cons), manual loading of data into HDFS
Day-2: Session-1: Big Data Ecosystem-Building Big Data ETL: The Universe of Big Data Tools-When to use which?
- Hadoop vs. Other NoSQL solutions
- Requirements for interactive, random access to data
- Hbase (column-oriented database) on top of Hadoop
- Random access to data with restrictions (max 1 PB)
- Suitable for logging, counting, time-series; less ideal for ad-hoc analytics
- Sqoop - Importing data from databases to Hive or HDFS (JDBC/ODBC access)
- Flume – Streaming data (e.g. log data) into HDFS
Day-2: Session-2: Big Data Management System
- Component movement, compute node start/fail :ZooKeeper - For configuration/coordination/naming services
- Complex pipeline/workflow: Oozie – managing workflows, dependencies, daisy chains
- Deployment, configuration, cluster management, upgrades, etc. (sys admin) :Ambari
- Cloud-based solutions : Whirr
Day-2: Session-3: Predictive analytics in Business Intelligence -1: Fundamental Techniques & Machine Learning Based BI :
- Introduction to Machine Learning
- Learning classification techniques
- Bayesian Prediction-preparing training files
- Support Vector Machine
- KNN p-Tree Algebra & vertical mining
- Neural Networks
- Big Data large variable problem -Random Forest (RF)
- Big Data Automation problem – Multi-model ensemble RF
- Automation through Soft10-M
- Text analytic tool-Treeminer
- Agile learning
- Agent-based learning
- Distributed learning
- Introduction to Open Source Tools for Predictive Analytics : R, Rapidminer, Mahut
Day-2: Session-4 Predictive Analytics Ecosystem-2: Common Predictive Analytic Problems in Government
- Insight analytics
- Visualization analytics
- Structured predictive analytics
- Unstructured predictive analytics
- Threat/fraud star/vendor profiling
- Recommendation Engines
- Pattern detection
- Rule/Scenario discovery –failure, fraud, optimization
- Root cause discovery
- Sentiment analysis
- CRM analytics
- Network analytics
- Text Analytics
- Technology Assisted Review
- Fraud analytics
- Real-Time Analytics
Day-3 : Sesion-1 : Real-Time and Scalable Analytics Over Hadoop
- Why common analytic algorithms fail in Hadoop/HDFS
- Apache Hama- for Bulk Synchronous distributed computing
- Apache SPARK- for cluster computing in real-time analytics
- CMU Graphics Lab2- Graph-based asynchronous approach to distributed computing
- KNN p-Algebra-based approach from Treeminer for reduced hardware operational costs
Day-3: Session-2: Tools for eDiscovery and Forensics
- eDiscovery over Big Data vs. Legacy data – comparison of cost and performance
- Predictive coding and technology-assisted review (TAR)
- Live demo of a TAR product (vMiner) to illustrate how TAR accelerates discovery
- Faster indexing through HDFS –velocity of data
- NLP or Natural Language Processing –various techniques and open-source products
- eDiscovery in foreign languages-technology for foreign language processing
Day-3 : Session 3: Big Data BI for Cyber Security – Comprehensive 360-Degree View from Data Collection to Threat Identification
- Understanding basics of security analytics-attack surface, security misconfiguration, host defenses
- Network infrastructure/ Large data pipe / Response ETL for real-time analytics
- Prescriptive vs predictive – Fixed rule-based vs auto-discovery of threat rules from Meta data
Day-3: Session 4: Big Data in USDA : Applications in Agriculture
- Introduction to IoT (Internet of Things) for agriculture-sensor-based Big Data and control
- Introduction to Satellite imaging and its application in agriculture
- Integrating sensor and image data for soil fertility, cultivation recommendation, and forecasting
- Agriculture insurance and Big Data
- Crop Loss forecasting
Day-4 : Session-1: Fraud Prevention BI from Big Data in Government-Fraud Analytics:
- Basic classification of Fraud analytics- rule-based vs predictive analytics
- Supervised vs unsupervised Machine Learning for Fraud pattern detection
- Vendor fraud/overcharging for projects
- Medicare and Medicaid fraud- fraud detection techniques for claim processing
- Travel reimbursement frauds
- IRS refund frauds
- Case studies and live demos will be presented wherever data is available.
Day-4 : Session-2: Social Media Analytics- Intelligence Gathering and Analysis
- Big Data ETL API for extracting social media data
- Text, image, meta data, and video
- Sentiment analysis from social media feed
- Contextual and non-contextual filtering of social media feed
- Social Media Dashboard to integrate diverse social media
- Automated profiling of social media profiles
- Live demos of each analytic will be conducted through the Treeminer Tool.
Day-4 : Session-3: Big Data Analytics in Image Processing and Video Feeds
- Image Storage techniques in Big Data- Storage solutions for data exceeding petabytes
- LTFS and LTO
- GPFS-LTFS (Layered storage solution for Big image data)
- Fundamentals of image analytics
- Object recognition
- Image segmentation
- Motion tracking
- 3-D image reconstruction
Day-4: Session-4: Big Data Applications in NIH:
- Emerging areas of Bio-informatics
- Meta-genomics and Big Data mining issues
- Big Data Predictive analytics for Pharmacogenomics, Metabolomics, and Proteomics
- Big Data in downstream Genomics processes
- Application of Big Data predictive analytics in Public health
Big Data Dashboard for Quick Accessibility and Display of Diverse Data :
- Integration of existing application platforms with Big Data Dashboards
- Big Data management
- Case Study of Big Data Dashboards: Tableau and Pentaho
- Using Big Data apps to push location-based services in Government
- Tracking systems and management
Day-5 : Session-1: Justifying Big Data BI Implementation Within an Organization:
- Defining ROI for Big Data implementation
- Case studies on saving Analyst Time for data collection and preparation –increase in productivity gain
- Case studies of revenue gain from saving licensed database costs
- Revenue gain from location-based services
- Savings from fraud prevention
- An integrated spreadsheet approach to calculate approximate expense vs. Revenue gain/savings from Big Data implementation.
Day-5 : Session-2: Step-by-Step Procedure to Replace Legacy Data Systems with Big Data Systems:
- Understanding a practical Big Data Migration Roadmap
- Essential information needed before architecting a Big Data implementation
- Methods for calculating volume, velocity, variety, and veracity of data
- Estimating data growth
- Case studies
Day-5: Session 4: Review of Big Data Vendors and their Products. Q&A session:
- Accenture
- APTEAN (Formerly CDC Software)
- Cisco Systems
- Cloudera
- Dell
- EMC
- GoodData Corporation
- Guavus
- Hitachi Data Systems
- Hortonworks
- HP
- IBM
- Informatica
- Intel
- Jaspersoft
- Microsoft
- MongoDB (Formerly 10Gen)
- MU Sigma
- Netapp
- Opera Solutions
- Oracle
- Pentaho
- Platfora
- Qliktech
- Quantum
- Rackspace
- Revolution Analytics
- Salesforce
- SAP
- SAS Institute
- Sisense
- Software AG/Terracotta
- Soft10 Automation
- Splunk
- Sqrrl
- Supermicro
- Tableau Software
- Teradata
- Think Big Analytics
- Tidemark Systems
- Treeminer
- VMware (Part of EMC)
Requirements
- Fundamental knowledge of business operations and data systems within the government domain
- Basic proficiency in SQL/Oracle or relational database concepts
- Basic understanding of statistics (at a spreadsheet proficiency level)
35 Hours
Testimonials (1)
The ability of the trainer to align the course with the requirements of the organization other than just providing the course for the sake of delivering it.