Berkay Şahin · MLOps Engineer

I keep your
pipelinesclustersinfrastructuredashboards
running in production.

Machine learning models are easy to train and hard to keep running. I build the pipelines that move them into production, the cloud they run on, and the monitoring that pages you before the user does.

notebook → productionlive
notebook
CI/CD
cluster
production
monitor
what i work on4 / 4 healthy
pipelines

Getting models from a notebook to a schedule

End-to-end ML pipelines on Databricks, deployed through GitLab CI/CD with Databricks Asset Bundles, authenticating as service principals over OAuth M2M rather than someone's personal token.

DatabricksGitLab CI/CDDABOAuth M2M
svc.01
infrastructure

The AWS account underneath

Provisioning and running the cloud resources ML workloads sit on, and keeping the bill from quietly growing past what the workload is worth.

AWSIaCCost review
svc.02
orchestration

Kubernetes, for workloads that finish

Cluster management for batch and training workloads — scheduling, lifecycle, and the container images they run from.

KubernetesDockerLinux
svc.03
observability

Knowing before the user tells you

Centralised logs through Promtail into Grafana Loki, metrics on Prometheus, and dashboards that answer the question you actually have at 3am.

GrafanaLokiPromtailPrometheus
svc.04
selected work

Hiring, or stuck on a pipeline?

I read everything that arrives and reply to anything that isn't a template.

Send a message →