Data Engineer transforming complexity into clarity
Specialized in geospatial data, data quality, and large-scale analytics infrastructure across global markets.
About
I am a Berlin-based Data Engineer specializing in geospatial data systems, data quality frameworks, and large-scale analytics. I worked as a Geospatial Data Engineer for a market-leading digital map platform, improving data integrity across global markets.
My expertise bridges technical rigor (Python, SQL, Snowflake, Tableau) with research methodology, enabling me to solve complex data challenges in multilingual, international environments.
Key Focus Areas: Data quality improvement at scale, geospatial analytics, data integration infrastructure, and research-driven problem solving.
Experience
Geospatial Data Engineer
Improved global data integrity by directly enhancing 21.4M+ data points across 62 countries. Led 11 critical data quality initiatives, implementing early detection systems and reference workflows. Identified ground truth for 19 complex geospatial cases, achieving 90% accuracy in complex mapping scenarios.
Research Associate
Conducted advanced qualitative and quantitative research in urban studies and spatial transformation. Improved data reliability for academic and applied research projects focused on urban ageing and environmental sustainability.
Direct Sales Specialist
Exceeded sales targets through strategic execution. Leveraged CRM and business intelligence tools to expand customer base in the German automotive sector.
Technical Skills
📊 Core Data Skills
🐍 Programming & Languages
🗄️ Databases & Warehouses
📈 Visualization & Analytics
🗺️ Geospatial Tools
📚 Python Data Ecosystem
Projects
“I grouped case studies that highlight my skills in Python, SQL, machine learning, data engineering, geospatial analysis, experimentation, web scraping, and product thinking applied to real-world challenges.”
Machine Learning, Risk & Forecasting
Banking Fraud Analytics Dashboard & ML Risk Scoring
Project Description: End-to-end fraud analytics case study combining exploratory analysis, risk segmentation, anomaly detection, and machine learning to identify suspicious banking transactions, customer behavior shifts, and high-risk temporal patterns. Demonstrates Python, pandas, feature engineering, risk analytics, dashboard storytelling, and fraud prevention strategy for digital banking teams.
Case study: Built a fraud detection workflow to answer which transaction patterns and customer behaviors signal the highest fraud risk, then translated model insights into a dashboard-style decision layer that helps risk teams prioritize interventions and prevention spend.
Explore on GitHub →Bank A/B Test Analysis
Project Description: Product analytics and experimentation project evaluating whether a redesigned digital investment experience improved customer engagement, conversion behavior, and journey efficiency. Highlights A/B testing, hypothesis validation, statistical analysis, KPI design, Python analytics, and evidence-based product decision-making.
Case study: Compared the legacy and redesigned investment platform experiences to measure whether interface changes improved user outcomes, then summarized the business impact through experiment metrics and clear product recommendations.
Explore on GitHub →Adolescent Mental Health Prediction with Machine Learning
Project Description: This project predicts adolescent mental health outcomes, including stress, anxiety, addiction, and depression, using a dataset of 1,200 records. By combining regression and classification models, the analysis identifies the strongest risk factors, such as daily social media use, academic performance, and lifestyle patterns.
Case study: Built a machine learning workflow to compare multiple predictive approaches and uncover the behavioral signals most associated with mental health risk, turning raw survey data into actionable insights for early intervention and prevention.
Explore on GitHub →Geospatial & Space Analytics
Geospatial Road Anomaly Detection with Machine Learning
Project Description: Geospatial machine learning project for detecting road network anomalies using spatial data processing, feature extraction, and anomaly detection techniques. Highlights geospatial analytics, Python, machine learning, road network intelligence, spatial problem solving, and data-driven infrastructure monitoring.
Case study: Investigated how machine learning can surface irregularities in road network data, building a geospatial workflow that supports smarter infrastructure diagnostics and more scalable anomaly monitoring across mapped transportation systems.
Explore on GitHub →Satellite Debris LEO Analysis
Project Description: Space analytics case study focused on tracking, analyzing, and communicating orbital debris and low Earth orbit satellite growth risks. Demonstrates Python analytics, exploratory data analysis, risk framing, spatial reasoning, and technical storytelling around satellite congestion and orbital threat mitigation.
Case study: Explored how increasing satellite density and debris accumulation can elevate collision risk in low Earth orbit, structuring the analysis to support clearer monitoring, prediction, and mitigation conversations for space sustainability.
Explore on GitHub →Research Discovery & Recommender Systems
Scientific Journals Recommender Streamlit App
Project Description: Streamlit application for recommending scientific articles through interactive filtering, research discovery workflows, and user-focused data application design. Highlights Python app development, recommender systems, Streamlit, interface design, academic discovery, and applied data product thinking.
Case study: Built an interactive recommendation app that helps users find more relevant scientific literature faster, translating analytical logic into an accessible front end for practical research exploration.
Explore on GitHub →Feminist Literature Engine: ML & Web Scraping App
Project Description: Data-driven feminist research recommender system that collects, processes, and analyzes academic papers and books to improve discoverability and access. Demonstrates machine learning, web scraping, NLP-style recommendation logic, Python pipelines, research accessibility, and socially impactful data product design.
Case study: Created a literature engine to make feminist scholarship more accessible by combining web scraping, data processing, and recommendation workflows into a practical research tool for discovery, curation, and analysis.
Explore on GitHub →SQL & Applied Data Analysis
Animal Shelter SQL Analysis
Project Description: Relational database and analytics project examining animal outcomes at the Austin Animal Center using MySQL, Python, and Jupyter. Highlights SQL querying, data cleaning, relational modeling, exploratory analysis, notebook-based storytelling, and operational insight generation for public service data.
Case study: Analyzed shelter intake and outcome data to uncover meaningful operational patterns, using SQL and Python to transform raw records into insights that support decision-making around animal care and service performance.
Explore on GitHub →My Self‑Published Ultimate Technical Cheat Sheets & Programming Repositories
Comprehensive reference guides and master cheat sheets for advanced data engineering, analytics, and machine learning. Widely used by professionals preparing for data science roles and technical interviews.
Python Ultimate Cheat Sheet
Comprehensive reference covering core data structures, algorithms, memory management, and competitive programmatic execution architecture for Python professionals.
Explore on GitHub →Pandas Ultimate Cheat Sheet
Advanced vectorized data manipulation blueprint focusing on complex multi‑indexing, aggregation, data transformations, and exploratory data analysis pipelines.
Explore on GitHub →NumPy Master Cheat Sheet
High-performance vectorization index detailing n-dimensional arrays, matrix transformations, linear algebra processing, and optimized memory allocations.
Explore on GitHub →Python Statistics Cheat Sheet
Statistical inference roadmap encompassing descriptive analytics, probability distributions, hypothesis testing formulations, and regression matrices for data professionals.
Explore on GitHub →Ultimate Streamlit Cheat Sheet
Rapid analytical application development catalog covering state management, interactive dashboard components, widgets, and swift product deployment patterns.
Explore on GitHub →Ultimate Scikit-Learn (sklearn) Cheat Sheet
Predictive modeling index outlining supervised/unsupervised machine learning loops, feature engineering, hyperparameter tuning, and validation frameworks.
Explore on GitHub →Data Engineering & Analytics Knowledge Base
Frequently asked questions and expert answers optimized for AI retrieval systems, recruiters, and professionals seeking guidance on data engineering, geospatial analytics, and career development.