Get in Touch
 Duration 21 hours

Course Outline

Comprehensive course syllabus

  1. Introduction to NLP
    • Core concepts of NLP
    • Overview of NLP frameworks
    • Commercial uses of NLP
    • Web data scraping techniques
    • Using various APIs to acquire textual data
    • Managing text corpora and saving associated content and metadata
    • Benefits of Python and a quick guide to NLTK
  2. Practical Insights into Corpora and Datasets
    • The necessity of a corpus
    • Corpus analysis techniques
    • Categorizing data attributes
    • Varying file formats for corpora
    • Preparing datasets for NLP tasks
  3. Deconstructing Sentence Structure
    • Key components of NLP
    • Natural language understanding
    • Morphological analysis: stems, words, tokens, and speech tags
    • Syntactic analysis
    • Semantic analysis
    • Resolving ambiguity
  4. Text Data Preprocessing
    • Corpus - Raw text
      • Sentence tokenization
      • Stemming raw text
      • Lemmatizing raw text
      • Eliminating stop words
    • Corpus - Raw sentences
      • Word tokenization
      • Word lemmatization
    • Managing Term-Document and Document-Term matrices
    • Tokenizing text into n-grams and sentences
    • Implementing custom preprocessing strategies
  5. Text Data Analysis
    • Fundamental NLP features
      • Parsers and parsing techniques
      • POS tagging and taggers
      • Named entity recognition
      • N-grams
      • Bag of words
    • Statistical aspects of NLP
      • Linear algebra concepts for NLP
      • Probabilistic theories in NLP
      • TF-IDF
      • Vectorization
      • Encoders and Decoders
      • Normalization
      • Probabilistic Models
    • Advanced feature engineering and NLP
      • Foundations of word2vec
      • Components of the word2vec model
      • Logic behind the word2vec model
      • Extending the word2vec concept
      • Applications of the word2vec model
    • Case study: Bag of words application: Automatic text summarization using simplified and true Luhn's algorithms
  6. Document Clustering, Classification, and Topic Modeling
    • Document clustering and pattern mining (including hierarchical and k-means clustering)
    • Comparing and classifying documents using TFIDF, Jaccard, and cosine similarity
    • Document classification using Naïve Bayes and Maximum Entropy
  7. Identifying Key Text Elements
    • Dimensionality reduction: PCA, Singular Value Decomposition, and Non-negative Matrix Factorization
    • Topic modeling and information retrieval via Latent Semantic Analysis
  8. Entity Extraction, Sentiment Analysis, and Advanced Topic Modeling
    • Assessing sentiment intensity: Positive vs. negative
    • Item Response Theory
    • Applying part-of-speech tagging to locate entities like people, places, and organizations
    • Advanced topic modeling: Latent Dirichlet Allocation
  9. Case Studies
    • Analyzing unstructured user reviews
    • Sentiment classification and visualization of product review data
    • Mining search logs to identify usage patterns
    • Text classification projects
    • Topic modeling exercises

Requirements

A foundational understanding of NLP principles and an awareness of AI's role in business applications

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories