Which sellers on a 110K-order Brazilian e-commerce marketplace are systematically causing delivery delays — after controlling for region and product category, so the analysis doesn't wrongly blame sellers shipping to remote areas or selling inherently slow-moving products.
881
sellers met the minimum order-volume threshold
371
exceeded their region + category-adjusted expected rate
$403K
modeled revenue-at-risk across flagged sellers
01SQLLoaded and joined raw order, seller, and product tables; built seller-level aggregation with a minimum-volume filter and region/category benchmark comparisons.
02PythonValidated the SQL output and modeled revenue-at-risk from the gap between expected and observed late rates.
03Power BIBuilt a joint state × category benchmark and an interactive dashboard for drilling into individual sellers.
METHODOLOGY NOTES — READ BEFORE DRAWING CONCLUSIONS
Expected late rate is a joint state × category benchmark — sellers are compared against others selling similar products into similar regions, not a flat average.
Revenue-at-risk is a modeled proxy (excess bad-review rate on late orders × late-order revenue), not a confirmed financial loss — the data has no field for actual platform interventions.
The analysis doesn't separate seller handling time from carrier transit time, so "late" reflects total delay to the customer, not seller fault in isolation.
A minimum of 20 orders is required before a seller enters the ranking, to avoid flagging low-volume sellers on noise alone.
IN PROGRESS
Project #2 — coming soon
Actively scoping the next analysis while applying to Business Analyst roles. Check back, or ask me directly what I'm working on.
TOOLS
QUERYING & DATA
SQL (MySQL), Excel, data cleaning & joins
ANALYSIS
Python (pandas), confound-adjusted benchmarking
VISUALIZATION
Power BI, DAX, Power Query
APPROACH
Stating a metric's limits alongside the metric itself