ENTERPRISE DATA PLATFORMS

Roshan
Dhamala.

Senior Data Engineer

From platform foundations
to production data & AI.

I help take enterprise lakehouses from proof of concept to production—setting integration standards, improving Spark performance, and building governed data and AI workflows.

Veeva Systems Since January 2022

ENTERPRISE ENGINEERING IMPACT

Selected outcomes.

More context

≈35%

Lower processing cost

Reduction in Spark-related processing cost.

≈99%

Pipeline reliability

Reported pipeline reliability.

≈40%

Higher data utilization

Increase in data utilization.

≈39.6M

Historical records

Historical records processed in a major ingestion initiative.

Approximate outcomes across my professional work; separate from the ongoing CRM migration.

SELECTED ENTERPRISE WORK

Ownership across the
data platform.

All work

ONGOING · ENTERPRISE INITIATIVE

Enterprise
CRM migration

From source-to-target mapping
to downstream reporting.

SANITIZED PROFESSIONAL SCOPE

DATA ENGINEERING TECHNICAL OWNERSHIP

Migration, integration, and reporting.

The initiative spans migration into the new CRM and the data platform that consumes it. My scope connects data quality and integration architecture with operational and analytical reporting.

  • Source-to-target mapping, reconciliation, and data quality
  • Integration architecture and ingestion from the new CRM
  • Downstream modeling, permissions, and governance
  • Coordination with business and technical stakeholders
Read the scope overview

PLATFORM ARCHITECTURE

POC → production lakehouse

Technical architecture for a Databricks platform with 100+ TB and 120+ datasets, with ingestion, modeling, and governance standards used across teams.

Platform overview

APPLIED DATA + AI

Documents → usable data

LLM/RAG workflows connecting document processing, structured extraction, metadata enrichment, and vector search with governed enterprise data.

Engineering scope

PUBLIC ENGINEERING PROJECT

DataNepal.

Inspectable code. Documented decisions.

A public data product that brings fragmented sources into a shared geographic model, with provenance and a static publishing boundary.

dltDuckDBdbtParquet

What the implementation demonstrates

  • Canonical identifiers and tested geographic crosswalks
  • Source, licence, vintage, and caveats enforced at export
  • Parquet / JSON delivery with browser-side queries
Inspect the repository

TECHNICAL PERSPECTIVE

Decisions, explained.

Writing

CURRENT ROLE

Senior Data Engineer

Veeva Systems · January 2025–Present

At Veeva since January 2022; promoted to Senior Data Engineer in January 2025. Lead architecture reviews and mentor engineers on Spark optimization, modeling, and production patterns.

View resume

CONTACT

Let’s talk data platforms.

Connect with me about enterprise data platforms, technical leadership, or applied AI.

Connect on LinkedIn