Question Bank
Data Engineering interview questions
A data pipeline must preserve the meaning of events despite duplicates, delays and failures. Discuss ingestion, transformation, storage and quality checks. At each step, explain how you would detect a problem and recover a correct result.
Showing 20 of 30 questions
What infrastructure have you used for model training: internal clusters or cloud services?
Have you built ML pipelines for training neural networks?
Do you have experience with FastAPI and building REST APIs?
Do you have experience with big data technologies such as Hadoop, Spark, or PySpark?
Given a CSV of exchange trades with timestamp and volume, find the time intervals with the highest trading volume. How would you determine the optimal window size?
A file of exchange trades does not fit in RAM. How would you find the interval with the highest volume in one pass?
A stream of exchange packets includes send and receive timestamps. How would you detect latency anomalies?
What data volumes and model training times have you worked with? Do you have experience with parallelism?
What experience do you have with Hadoop, Spark, and MapReduce?
Describe your experience with PySpark and distributed computing. How does MapReduce work?
Какие метрики бы предложили для оценки дрифта данных? И для вторичной валидации моделей каждый раз бы их использовали?
И какие риски мы можем переобучить на каких-то форты данных?
Под что может больше строковые базы данных пригодиться?
А уверенность, это что за такая штука такая? Как ее добыть из данных?
а есть какая-то структура данных, которая может, в принципе, тоже достаточно быстро искать, но при этом в отсортированном виде быстро получать последовательность?
А у вас на счет данных получается как-то Бигдад используется?
На какие виды можно разделить типа данных в питании?
Prepare for your next interview with Vibe Interview.
Download Vibe Interview