Skip to content

Data Engineer · AI Engineer · Builder

DorYardeni

From raw events to AI agents - built end to end.

01 / noise

Raw. Messy. Everywhere.

Large-scale event data landing from Kafka, Event Hubs, S3 and half a dozen databases. Nobody trusts it yet, and nobody owns it. I start here, usually with the analysts and product people who need something out of it.

  • Kafka
  • Event Hubs
  • S3
  • PostgreSQL
  • Elasticsearch

02 / rhythm

Structure out of noise.

A live streaming ETL pipeline on Spark and Databricks: 10,000 events a second at peak, hundreds of terabytes a month, Delta Lake underneath, Terraform and CI/CD keeping it honest. I built it, and I'm who the team asks when something in it breaks.

  • Databricks
  • Apache Spark
  • Delta Lake
  • Structured Streaming
  • Terraform

03 / harmony

Then it learns.

An AI agent that works like an analyst: it understands company-specific domain data and runs real analysis workflows - about 1000× faster than a human analyst, at 1/100 of the cost. It wasn't on anyone's roadmap. I scoped it, built it, and pitched it until it was.

  • LLMs
  • AI Agents
  • RAG
  • LangGraph
  • MLflow

04 / performance

Shipped. Owned. End to end.

Frontend, backend, security and evaluation - from concept to a product used internally and pitched to customers. I wrote the code and ran the conversations with product, sales and security that turned it into an offer we win with.

  • TypeScript
  • Node.js
  • React
  • Vue
  • AWS Lambda
  • Azure Functions

Experience

2022 - today

  1. 2025 - 2026

    1000× faster · 1/100 the cost

    AI Engineer

    CYMOTIVE Technologies

    LLMsLangChainLangGraphRAGMLflowPython

    Initiated and independently built a new AI product from concept to proof-of-concept: an AI agent that works as an analyst over company-specific domain data - about 1000× faster than a human analyst at roughly 1/100 of the cost.

    • Designed and implemented an AI agent that understands domain data and runs real analysis workflows end to end.
    • Built the full product flow: frontend, backend, security, model evaluation, agent logic and integration with internal data sources.
    • Made the company materially more efficient in time and money, and made our commercial offers far more competitive than competitors'.
    • Presented and drove the initiative internally until it became both an internal productivity tool and a customer-facing product candidate.
  2. 2023 - 2025

    10k events/s · 100s of TB/month

    Data Engineer

    CYMOTIVE Technologies

    DatabricksApache SparkDelta LakeKafkaTerraformSQL

    Built a live streaming, end-to-end ETL pipeline that consumes 10,000 events per second at peak - hundreds of terabytes every month - on Spark and Databricks, and led a full system migration from Elasticsearch into Databricks.

    • Built live ETL and streaming pipelines with Databricks, Spark Structured Streaming, Delta Lake and Kafka.
    • Created, optimized and maintained database architectures, schemas and complex SQL queries over large-scale event data.
    • Automated cloud data infrastructure with Terraform and CI/CD.
    • Go-to expert in the team for data engineering, Databricks, Spark and secure coding practices.
  3. 2022 - 2023

    End-to-end features

    Full Stack Engineer

    CYMOTIVE Technologies

    TypeScriptNode.jsVueReactPostgreSQLAWSAzure

    Developed backend services, REST APIs and data-driven applications in TypeScript and Node.js, with Vue and React frontends and serverless functions on AWS and Azure.

    • Designed and implemented product features from frontend interfaces to backend APIs and serverless functions.
    • Worked with AWS Lambda, Azure Functions and cloud services across AWS and Azure.
    • Implemented database schemas, SQL queries and backend integrations with PostgreSQL.
    • Contributed to CI/CD pipelines, infrastructure automation and deployment workflows.

Things I build
on my own time

  • Jagura

    2025

    An SQL interface for managing containers. Jagura behaves like any SQL database, plus a CONTAINER data type: start, stop, restart, pause or kill containers, read their metadata and run commands inside them - all from ordinary SELECT statements.

    • Docker
    • SQL
    • Node.js
    • TypeScript
  • An AI-based competitive intelligence platform. It tracks competitors, turns their product, pricing and security changes into insights, compares them side by side and ranks everything by impact automatically - with every claim traced back to its source.

    • AI Agents
    • LLMs
    • RAG
    • Context Engineering
    • Prompt Engineering
  • One agent skill per engineering book I've read - each distilling the book's rules and methodology into a single SKILL.md, so I can call a book into my agent while developing. Covers AI engineering and RAG, data-intensive systems, refactoring, pragmatic engineering and UX. Works in Cursor, Claude Code and anything else that speaks the open Agent Skills format.

    • Agent Skills
    • LLMs
    • Context Engineering
    • Cursor
    • Claude Code
  • An article I wrote after benchmarking PySpark against Scala: identical performance on the native DataFrame and SQL APIs, but up to 10× faster in Scala once heavy UDFs enter the picture - and the rewrite that finally stopped a streaming job from running its driver out of memory.

    • Writing
    • Apache Spark
    • PySpark
    • Scala
    • Benchmarking

About

From raw events to AI agents - built end to end.

Data and AI Engineer with a strong backend, data engineering and cloud infrastructure background. I've worked the whole product lifecycle - frontend, REST APIs, large-scale Spark pipelines, cloud infrastructure, AI agents, RAG and context engineering - and I like taking a proof-of-concept all the way to a product people actually use.

ENGLISH EXCELLENT · HEBREW NATIVE

1000×
faster than a human analystAI agent
1/100
the cost of a human analystAI agent
10k/s
events ingested at peakstreaming ETL
100s TB
of data processed monthlystreaming ETL
Languages
TypeScriptPythonSQLJavaScriptScala
AI Engineering
LLMsAI AgentsRAGContext EngineeringPrompt EngineeringLangChainLangGraphMLflow
Data & Big Data
Apache SparkDatabricksDelta LakeStructured StreamingKafkaETLData Modeling
Backend & Frontend
Node.jsREST APIsVueReact
Cloud
AWS S3LambdaEC2SQSAzure Event HubsBlob StorageAzure Functions
Infra & DevOps
TerraformGitHub ActionsDockerCI/CD
Databases
DatabricksPostgreSQLElasticsearchSQL optimizationSchema design
Soft Skills
Product ownershipCross-team communicationPitching & demosMentoringTechnical writing

Contact

Let's build
something.

01Particles02Music