Video summary

This role Pays more than Software Engineers | TechWithPrateek Kartikhustles | Data Engineer Roadmap

Main summary

Key takeaways

Technology

Summary

Senior data engineer Prateek Singhal discusses what data engineering involves, how it differs from software engineering, hiring and compensation, interview preparation, and career paths for people entering the field.

What Data Engineers Do—and Why the Work Is Challenging

  • Data engineering involves writing code to ingest, transform, and deliver data through pipelines. The data itself adds complexity: a pipeline that works today can fail if an upstream source changes its schema, volume, or delivery timing.
  • Large datasets are difficult to understand without examining patterns over time and knowing the business context. Data engineers need programming, data, and business skills.
  • Pipeline work also includes troubleshooting, optimizing processing, and backfilling or correcting older data after failures.
  • Singhal cautions against assuming data engineering is an easier, less-coding alternative to software engineering. He describes it as demanding and multi-layered.

Demand, Pay, and AI

  • Singhal sees strong hiring demand for data engineers, partly because AI applications depend on clean, structured, and reliable data. He expects the need for data engineering to grow, while noting that there are fewer data engineer roles than general software developer roles.
  • Based on companies he has interviewed with or worked at, he says data engineer compensation is generally comparable to software engineer compensation. Data analysts typically have a lower compensation ceiling. He suggests supply and demand could shift salaries over time, but does not claim data engineers universally earn more.
  • AI tools such as Claude Code and Copilot can help write pipeline code, analyze data, generate SQL, and revise queries. Their usefulness depends on an engineer’s ability to guide and verify them; otherwise, they may waste compute time or produce unsuitable results.
  • Singhal argues that AI can increase the amount of code produced without proportionally increasing production-ready code, adding to the burden of human review.

How AI Is Being Applied to Data Work

  1. Self-service analytics: AI systems can interpret a business question, use business and table context to generate and run SQL, check results, and refine the answer iteratively. This aims to reduce the time spent relaying questions between business teams, product managers, and analysts. Singhal says real-world data complexity still makes this difficult.
  2. Self-healing pipelines: AI could identify a failed pipeline, detect changes in incoming data, propose code and tests to address the failure, and prepare the fix for human review. Singhal describes this as an emerging approach rather than something widely deployed in production.

Skills and Interview Preparation

Singhal recommends building a foundation in the following areas:

  • Python and data structures and algorithms (DSA): Python is common, though Scala or Java may also be used. Product-company interviews may include coding rounds comparable to software engineering interviews.
  • SQL: Candidates should be comfortable with aggregations, window functions, common table expressions (CTEs), and writing queries to answer analytical questions—not just basic SQL commands.
  • Data modeling: Star schemas and Kimball-style modeling are relevant. Senior interviews may test how candidates model data for a business scenario; Singhal says this can be decisive at the senior level.
  • Data system design: Be prepared to discuss batch and streaming processing, real-time requirements, storage and compute choices, partitioning, performance, and Lambda or Kappa architectures.
  • Spark: Spark and PySpark are widely used for big-data processing, and Spark optimization may be tested. Some employers may also ask about platforms such as Databricks, cloud services, or AWS S3.
  • Database and cloud fundamentals: Understand how SQL databases store and load data, why their design differs from big-data systems, and how cloud object storage works.
  • Behavioral and scenario questions: Interviews may cover project decisions, collaboration, stakeholder alignment, and company-specific leadership principles. Practical scenarios may ask how to ingest, transform, and process data arriving on a schedule.

Routes into Data Engineering

Singhal outlines three possible paths:

  1. Start in data analytics: Build strong SQL and Python skills, develop a solid understanding of data, then learn Spark and other engineering skills while working.
  2. Start in software engineering: Use existing programming, systems, deployment, and CI/CD experience as a foundation, then specialize in data.
  3. Target data engineering directly: Develop software-engineering-level Python and DSA skills, strong SQL and database fundamentals, cloud knowledge, and introductory data-modeling skills.

Singhal notes that college projects often lack the scale of production data systems: a few hundred megabytes of sample data does not reproduce the challenges of processing terabytes daily. His advice is to focus on fundamentals, then gain large-scale experience on the job.

Learning Resources and Professional Profile Advice

  • For programming foundations, Singhal recommends Harvard CS50.
  • For SQL practice, he mentions HackerRank, LeetCode, W3Schools SQL, and SQLBolt. HackerRank and LeetCode can also be used for Python and DSA practice.
  • He says data engineering learning resources are fragmented, with no single resource covering the whole field. He suggests searching separately for areas such as data modeling and system design.
  • For LinkedIn, he recommends a professional photo and cover image, a clear and accurate headline, and evidence of completed work such as projects, published research, or exam-based certifications. He advises avoiding inflated labels and treating LinkedIn as a professional network rather than a personal social feed.
  • A profile alone does not secure a job: applications, referrals, a resume, interview preparation, and luck can all matter.

Career Advice

Singhal’s main lessons are to build a consistent learning habit—around one or two hours a day—and to develop communication and business understanding alongside technical skills. Engineers work with people and stakeholders, not just tools. Understanding what users and business teams need can lead to better products and stronger career growth.

Main Speakers and Sources

  • Kartik Agarwal — Host; introduces the discussion and asks about the role, hiring, preparation, and career prospects.
  • Prateek Singhal — Guest; senior data engineer who discusses his experience at Microsoft and Samsung and provides technical and career guidance.

Rate this summary

Your feedback will help improve summaries.

Improve this summary

Reprocess with a stronger model when the summary feels incomplete or inaccurate.

Pro

Translate summary in another language

Pro

Ask questions to this video

Chat for follow-up questions, clarifications, and source-backed answers.

Coming soon

Share this summary

Original video