Video summary

Mata Kuliah Humaniora Digital FIB UI : Pengantar Coding Humaniora Digital

Main summary

Key takeaways

Educational

Summary

The lecture introduces Python coding for processing text in the digital humanities. Digital text collections, or corpora, often need to be cleaned and structured before analysis. A corpus may be incomplete or damaged, contain unreadable or irrelevant material, or arrive in an unstructured form.

The session demonstrates basic text-processing operations and introduces Google Colab and Jupyter notebooks as tools for writing, running, and documenting code.

Main Concepts

  • Why process text: Prepare corpus data so it is more consistent and suitable for analysis.
  • Common text-processing tasks:
    • Case folding: Convert text to lowercase so differently capitalized forms are treated consistently.
    • Punctuation removal: Strip punctuation that is not needed for an analysis.
    • Tokenization: Divide text into smaller units, usually words.
    • Stopword removal: Exclude common words that may not be meaningful for a particular analysis.
  • Why use code: Coding offers flexibility and customization beyond what may be available through graphical-interface tools. Programming languages have formal syntax that must be followed so a computer can interpret instructions.
  • Notebooks: A Jupyter notebook (.ipynb) combines executable Python code with written explanations. This makes it useful for documenting analysis and sharing work for collaboration.
  • Notebook cells: Code cells contain Python instructions; text cells contain explanations and are written using Markdown.

Practical Workflow and Instructions

1. Open and Set Up Google Colab

  1. Open Google Colab in a browser and log in to a Google account.
  2. Select New notebook.
  3. Close the release-notes sidebar if it appears.
  4. Rename the notebook—for example, Basic Coding Digital Humanities. Colab saves changes automatically.
  5. Remove the initial blank cell if desired, then add a text cell for the notebook’s title and introductory notes.

2. Document Code as You Work

For each exercise:

  1. Add a new code cell and write the Python code.
  2. Run the cell using its play button.
  3. Check the output beneath the code.
  4. Add a text cell below it to explain what the code does and what output it produces.
  5. If an error occurs, check the syntax, quotation marks, and indentation. Python code that uses indentation should be indented consistently.

The instructor emphasizes that a notebook should record both the code and the reasoning or explanation behind it—not just the code itself.

3. Learn Basic Python Data Structures and Output

  • Variables store data; the instructor compares them to boxes.
  • Strings are text values enclosed in matching single or double quotation marks.
  • Numbers are written without quotation marks.
  • Lists hold multiple items in order, using square brackets. Their positions are zero-indexed: the first item is at index 0.
  • Dictionaries hold labeled values, using curly braces. Items are accessed by their labels rather than by position.
  • The print() function displays text, values, and variable contents.
  • The + operator can join strings (concatenation).
  • The examples show how to retrieve list items by index and dictionary items by label.

4. Use Imports and Conditional Logic

The lecture demonstrates a date-based example:

  1. Import Python’s date/time functionality and give it a shorter alias.
  2. Create date variables, including a date written in ISO format.
  3. Calculate a relative date, such as the following day.
  4. Use an if statement to compare dates and display different messages depending on whether a condition is met.

This illustrates how program logic can make different decisions based on data.

5. Clean and Transform Text

  • Change text to lowercase: Use a lowercase operation and print the result.
  • Change text to uppercase: Convert text to lowercase first if needed, then apply an uppercase operation.
  • Remove punctuation with regular expressions:
    1. Import Python’s regular-expression module.
    2. Use a substitution operation to find non-alphanumeric characters.
    3. Replace matches with an empty string and display the cleaned text.
  • Correct repeated spelling errors: Use a regular-expression pattern to find a recurring misspelling and replace it with the intended spelling. The instructor presents this as an efficient way to handle many instances.

6. Tokenize Text and Remove Stopwords

  1. Use split() to divide text into tokens, typically at spaces, so each word becomes an item in a list.
  2. Treat the resulting list as input for later analysis.
  3. Define a collection of Indonesian stopwords—for example, conjunctions that are not useful for the intended analysis.
  4. Create a cleaned token list that retains words from the original tokens except those included in the stopword collection.
  5. Run the code and document the transformation in a text cell.

Speakers and Sources Featured

  • Speaker: One instructor or lecturer, unnamed in the subtitles. No other speakers are identifiable.
  • Tools and concepts discussed: Python, Google Colab, Jupyter Notebook, Markdown, regular expressions, and text-processing operations.

Rate this summary

Your feedback will help improve summaries.

Improve this summary

Reprocess with a stronger model when the summary feels incomplete or inaccurate.

Pro

Translate summary in another language

Pro

Ask questions to this video

Chat for follow-up questions, clarifications, and source-backed answers.

Coming soon

Share this summary

Original video