Hey, I'm

Tushar

Data Scientist · Data Engineer · Analytics Engineer

$142K platform spend brought under governance at Barton Malow
211 automated data quality checks shipped to production
10M+ records validated across telecom billing pipelines at Oracle
90× faster billing automation built at Accenture
Top 10 of 60+ teams, Penn State Nittany AI Challenge 2026

$ git log --oneline --graph

How I got here

Five years of watching production billing pipelines fail quietly set the direction. Every stop since has been about building the tools that catch a failure before it becomes an outage, not just cleaning up after one.

7f3a2c1 Accenture · Jun 2019 – Feb 2022

fix: automate batch bill-run workflows, 1 day → 2 hours

First real lesson in production data: a 4.5M-customer billing base doesn't forgive manual steps. Resolved 50+ recurring data quality issues and cut processing time 90×.

c9e81b4 Oracle Corporation · Feb 2022 – Dec 2024

feat: validate telecom billing at scale, 600K+ subscribers

Moved from fixing issues after the fact to designing the checks that catch them first — reconciliation reports and anomaly rules holding 98%+ accuracy within SLA. Earned "Best Upcoming Talent," FY23 Q3.

a10d5e7 Penn State Research · Sep 2025 – May 2026

feat: forecasting + agentic auditing, 92% accuracy

Went back to build the theory behind the practice: statistical modeling, and a first prototype of autonomous agents doing continuous compliance auditing.

f42b901 Barton Malow · May 2026 – Aug 2026

feat: cost analytics + DQ framework shipped to production

Where the two threads met: $142K in platform spend governed, 211 automated quality checks live across 23 tables in a $159B CRM pipeline.

The throughline: every role added a layer — fixing, then preventing, then modeling, then governing. The Barton Malow work is what all four look like combined into one pipeline.

$ cat PRINCIPLES.md

The pragmatic builder

Everything I build is grounded in a production mindset.

"Understanding existing systems deeply before changing them, designing for the next person who has to maintain my work, and communicating technical findings in a way non-technical stakeholders can act on." That’s the standard I actually build to.

$ cat experience.log

5+ years turning messy real-world data into decisions leadership can act on

Data Science Intern

Barton Malow Featured

May 2026 – Aug 2026 Michigan, United States
  • Built a Databricks Cost Analytics pipeline and 4-tab leadership dashboard from Unity Catalog billing/compute data, surfacing $142K in lifetime platform spend and $13.2K in idle resource cost across 885 unused objects. Automated an ownership-resolution hierarchy across 933 resources that cut the unowned-spend gap from 85.6% to 54%.
  • Designed a Claude AI cost-attribution dashboard isolating agent query costs from shared Databricks serverless compute (20+ users, 8,000+ queries, $0.53 avg cost/query), revealing AI's share of warehouse spend grew 5.5× in 5 months (1.5% → 8.3%), giving leadership real-time cost visibility before it became a budget concern.
  • Engineered a 211-rule automated data quality framework spanning 23 production tables in a $159B CRM pipeline, achieving 100% clean execution while using DQX's AI-driven profiler to accelerate future rule generation.
DatabricksPySparkSQLPythonClaude CodeGitHub Actions

Research Assistant, Data Science

Pennsylvania State University

Sep 2025 – May 2026 Pennsylvania, United States
  • Developed forecasting models on historical attendance data, achieving 92% accuracy and enabling data-driven vendor allocation and capacity planning.
  • Reduced event planning timeline 83% (12 months → 2 months) by building data-driven workflows, dashboards, and scenario simulations that aligned stakeholders on constraints and priorities.
  • Prototyped autonomous AI agents for continuous auditing and ethical compliance, including WORM-based audit trails, ontology-driven fairness checks, and cybersecurity safeguards aligned with ISO/IEC 42001.
Statistical ModelingForecastingAgentic AIPythonSQL

Staff Data Consultant

Oracle Corporation

Feb 2022 – Dec 2024 Bangalore, India
  • Validated end-to-end billing data pipelines for 600K+ telecom subscribers using SQL-driven quality checks, anomaly detection rules, and reconciliation reports, maintaining 98%+ accuracy within SLAs.
  • Built and automated KPI dashboards (churn, usage patterns) with 99% on-time delivery, enabling leadership to track performance in weekly reviews.
  • Led onsite test efforts, coordinating with offshore analytics teams to deliver readiness dashboards supporting a successful, on-time go-live.
Oracle DatabaseSQLData AnalysisBusiness Analysis
Award Oracle "Best Upcoming Talent," FY23 Q3

Data Consultant

Accenture Solutions Pvt Ltd

Jun 2019 – Feb 2022 Pune, India
  • Identified and resolved 50+ recurring data quality issues in production billing datasets for a 4.5M-customer base, improving billing accuracy and customer trust.
  • Automated batch bill-run workflows using SQL scripts and scheduling, cutting processing time from 1 day to 2 hours (90× faster) and manual effort by 92%.
  • Designed a SQL-based data validation framework (row-level checks, reconciliation, exception flags) that reduced discrepancies by 60% and shortened month-end close by 10%.
SQLShell ScriptingExcelClient Communication
Award Accenture "Shared Success Catalyst," Addressing Client & Community Needs

$ ls ./projects

Real systems, real numbers

AI Sentinel: Compliance Monitoring

Autonomous compliance agents detecting AI-driven regulatory risk in real time. Top 10 of 60+ teams, Penn State Nittany AI Challenge 2026. 5 detection rules mapped to CMS federal regulations (42 CFR §483).

PythonStreamlitOllamaGemini
View repo →

Resume–JD Matching via Deep Learning

3-stage semantic ranking pipeline: TF-IDF baseline → SBERT embeddings → supervised neural refinement into a calibrated match probability. 120 resumes ranked against 2,277 job descriptions.

PythonPyTorchSBERT
View repo →

Worldwide LLM Sentiment Analysis

End-to-end NLP + geospatial pipeline tracking global sentiment toward LLMs across 305K+ tweets. Location-normalization cut unresolved geography from 72.8% to 43.85%.

PythonNLTK/VADERPlotly
View repo →
211checks
23tables
84%records flagged

Databricks Data Quality Framework

Automated data quality monitoring built on Databricks DQX. Surfaced that 84% of active records carried at least one issue, and root-caused a ~39% completeness gap in a core reporting table.

CI: lint + schema validation Soft-fail: never breaks the upstream job
DatabricksPySparkSQLPython
View repo →
$142Kspend governed
54%unowned spend (from 85.6%)
error caught pre-ship

Databricks Cost Analytics & FinOps

FinOps cost-attribution pipeline on Unity Catalog system tables. Cut unowned spend from 85.6% to 54% and caught a 2x measurement error in the AI-cost methodology before it shipped.

CI: lint + job-config validation Idempotent: 0 duplicate rows on every rerun
DatabricksPySparkSQLStar Schema
View repo →

Scholar Swipe: Scholarship Scraper

Automated scraping pipeline extracting scholarship listings into a centralized PostgreSQL (Supabase) database, with upsert-based dedup and automated summarization.

PythonBeautifulSoup4PostgreSQL
View repo →

Alzheimer's Market Overview (US)

TRx-based pharmaceutical market analysis across brands, specialties, and channels. 202,499-row dataset, built as parallel Python (Plotly) and Tableau dashboards for direct platform comparison.

PythonPandasPlotlyTableau
View repo →

Penn State Alumni Event Forecasting

Multi-year (2022–2025) event performance analysis with scenario-based 2026 forecasting, projecting 235 attendees and $49/attendee revenue via a blended trend + recency-weighted model.

PythonPandasForecasting
View repo →

Deep Dive · Barton Malow, 2026

Databricks DQX (Data Quality Extended) Implementation

Problem

23 production tables in a $159B CRM pipeline had no automated way to catch bad data before it reached downstream reports.

Why

Bad records silently corrupt reporting and decisions. The earlier one is caught, the cheaper it is to fix.

How

211 automated rules run on a scheduled bronze/gold pipeline, CI-gated and soft-fail, routing findings to record owners every week.

Active records carrying at least one data quality issue

84%
211automated checks in production
23production tables covered
~39%completeness gap root-caused
100%clean execution across the $159B pipeline

Deep Dive · Barton Malow, 2026

Databricks Cost Analytics pipeline

Problem

85.6% of Databricks platform spend had no identifiable owner, and AI agent costs were invisible inside shared compute.

Why

You can't govern spend you can't attribute, and AI usage was growing too fast to stay unmeasured.

How

An automated ownership-resolution pipeline across 933 resources, plus a dedicated AI cost-attribution layer, feeding a 4-tab leadership dashboard.

Unowned spend, after ownership resolution

85.6% 54%

AI's share of Databricks warehouse spend, over 5 months

1.5% 8.3%
$142Klifetime spend surfaced
$13.2Kidle cost, 885 unused objects
measurement error caught before it shipped

$ cat research.md

Published work and competitive recognition

Published Paper · Springer Nature

Autonomous Multi-Agent Governance: AI Sentinel Framework for Mitigating Psychological Restraints in AI-Driven Fall Management

Accepted and presented at SEET 2026 (2nd International Conference on Software Engineering of Emerging Technologies), Penn State Behrend, PhD Research Track. Listed on page 76 of the conference abstract book. To be published in the official conference proceedings by Springer Nature.

View conference →
Competition · Penn State Nittany AI Challenge 2026

AI Sentinel: Top 10 of 60+ Teams

Won the Prototype Phase and MVP Phase, advanced through funded phases to a Top 10 finish at the Pitch Contest, Hintz Family Alumni Center, University Park.

Award · Oracle Corporation

"Best Upcoming Talent" (FY23 Q3)

Recognized for exceptional contributions and high-impact performance on the Telia Company billing analytics engagement.

Leadership · Penn State University

Student Senator

Representing students in discussions with university leadership; championing AI literacy and equitable AI access across Penn State's Commonwealth Caucus.

$ cat skills.yaml

Tools I actually ship with

Programming

Python (Pandas, NumPy, Matplotlib, Seaborn, scikit-learn, TensorFlow, PyTorch, Keras) · SQL · R · Java

Data Engineering & Cloud

Databricks · PySpark · Unity Catalog · Azure · AWS · Oracle Database · PostgreSQL · PL/SQL

ML & Statistics

Regression · Classification · Clustering · NLP · Time Series Forecasting · Feature Engineering · Model Evaluation (AUC, F1)

BI & Visualization

Tableau · Power BI · Plotly · Matplotlib · Seaborn

Dev Tools

Git · GitHub Actions · Jupyter · Jira

AI Tools

Claude Code · Claude Cowork · Claude Skills · Claude Projects · Claude MCP · ChatGPT · GitHub Copilot · Cursor IDE · Perplexity · Ollama · Google Gemini · Julius AI

$ cat education.log

Background

Pennsylvania State University

Aug 2025 – Dec 2026

Master of Data Analytics · GPA 3.96/4

Statistical Analysis, Data-Driven Decision Making, Predictive Analytics, Deep Learning, Natural Language Processing.

KIIT University

May 2015 – May 2019

B.Tech, Electronics & Instrumentation Engineering · GPA 3.3/4

Control Systems, Digital Electronics, OOP, Data Structures & Algorithms, Web Technology, Artificial Intelligence.

$ mail tushar

Let's build something that ships.

Open to full-time Data Scientist / Data Engineer / Analytics Engineer roles starting Dec 2026.