Skip to main content

Automating Metadata Lineage with Revenue Analytics Software

We helped a revenue analytics software platform overcome manual monthly data integration, lineage gaps, and pipeline errors. Through custom enterprise software development, we built a scalable unified data layer with automated lineage import, an AI agent, and an Angular frontend.

Digital publisher platform
Banner

Client

The client is an advertising data platform that provides revenue analytics for publishers. The publisher platform brings together ad data from both direct and programmatic channels, helping publishers analyze placement performance and identify billing irregularities.

With the increasing number of data sources plugged into the system, manual configuration and poor metadata documentation have become the main roadblocks.

Business Challenge

The data team managed an integration platform that scaled in sources and users faster than its infrastructure could support, creating challenges across three main areas.

Operational Pain

Integrating direct and programmatic ad sources required manual, repetitive setup for every data stream entering the BI layer. Dependency conflicts between internal apps caused frequent pipeline failures, forcing engineers to debug instead of develop.

Technical Bottlenecks

The frontend lacked a unified architecture across dashboards, reports, and connections. Without CI/CD for the internal CLI utility and zero metadata creation automation, adding a new data source heavily increased the documentation burden.

Business & Scaling Risks

Onboarding new data engineers took weeks without standardized tooling, and every BI integration required custom code. Manual metadata imports into Atlan couldn’t keep pace with source growth, putting revenue reporting accuracy at risk.

Solution

Jelvix addressed the full revenue analytics software stack through five parallel workstreams, drawing on big data analytics tools and AI development to map each one to a specific layer of the business challenge.


Component 1: Internal CLI Tool with CI/CD Distribution

An internal CLI tool standardizes repetitive data engineering tasks across all teams. The internal CLI is automatically distributed via GitLab CI/CD and Mise; thus, updates get pushed to all the teams without manual distribution. Documentation allows engineers to self-onboard with adoption assistance during distribution.

Component 2: Data Platform Stabilization

Dependencies that prevented the pipeline from running have been removed, and automatic startup has replaced the repetitive manual configuration performed by engineers each day. The stable platform uses Python, Apache Airflow, and PostgreSQL.

Component 3: Data Lineage Automation

Relationships between lineage across data sources, transformations, and dashboards can be automatically discovered through OpenMetadata, without any kind of manual effort. Then an AI agent processes the relationships and imports those as files to Atlan. The engineers then review the auto-generated source, table, and column descriptions before final ingestion.

Component 4: Unified Frontend

The front-end was redeveloped using TypeScript and Angular in four modules – dashboards, reports, data sets, and connections, which use Highcharts for data visualization purposes. There is now one common architecture instead of four modules.

Component 5: Revenue Analytics QA

Given that billing accuracy is critical to building trust among publishers, test automation was spearheaded by two QA teams in the revenue pipeline. The use of GenAI in test automation processes ensured test coverage without increasing the size of the QA team in direct proportion to the expansion of the platform’s metadata management automation.

  • 5

    Core Solution Workstreams

  • AI

    Automated Lineage Engine

  • 4

    Unified Frontend Modules

  • Business Architecture
  • Team
  • Development in Detail
  • Technology Stack
  • The publisher platform was rebuilt across five domains, each addressing a specific layer of the revenue analytics stack.

    Data Ingestion Domain: Apache Airflow, Python, and PostgreSQL help ingest data directly from advertising sources and programmatic advertisements into a single place. This domain operates smoothly because dependency issues have been resolved and auto-startup has been configured.

    Metadata and Lineage Domain: OpenMetadata and Atlan, combined with the AI agent for import file generation, handle automated documentation of sources, tables, columns, and lineage relationships. This domain handles data lineage automation and eliminates the manual management bottleneck. 

    Internal Tooling Domain: This custom CLI utility, delivered through GitLab CI/CD and Mise, enables standardized workflows for data engineering across the organization. Updates to the tool are automatic, and the documentation portal enables self-onboarding for new engineers.

    Frontend and Visualization Domain: TypeScript, Angular, and Highcharts enable BI integration through a single interface for dashboards, reports, datasets, and connections. The consistent architecture in all four applications provides data engineers with an integrated workspace for revenue analytics.

    Quality and Reliability Domain: Testing frameworks and GenAI-powered testing practices provide coverage of the entire revenue analytics pipeline. The deployment process of revenue-critical components is based on Kubernetes and Docker. 

  • The project required data engineering depth, metadata automation expertise, and full-stack frontend delivery:

    Data Engineer: Data platform stabilization, dependency resolution, automated lineage import into OpenMetadata

    CLI Tool Developer: Core tool functionality, GitLab CI/CD integration, internal documentation, AI agent for Atlan import files

    Frontend Engineers (2): Angular development across Dashboards, Reports, Data Sets, and Connections, Highcharts integration

    Software Developer in Test: Leadership across two QA teams, GenAI-assisted test automation, revenue analytics pipeline QA

    DevOps Engineer: Kubernetes, Docker, CI/CD pipeline maintenance

  • Five workstreams ran in parallel to move the platform from a fragmented, manually managed state to a fully automated data infrastructure.

    1. Platform stabilization
    A phase of IT consulting determined root causes before implementing fixes, resulting in the resolution of dependency conflicts and the automatic configuration of startup. The Python, Apache Airflow, and PostgreSQL stack was stabilized before metadata and tooling work started.

    2. Internal CLI tool development
    The core functionalities were developed initially, followed by the integration process of GitLab CI/CD and Mise distribution. The internal documentation was made available to the self-service docs platform.

    3. Metadata automation
    Lineage import through automation in OpenMetadata established relationships between sources, transformations, and dashboards. An AI agent was next built for creating Atlan import files for sources, tables, columns, and lineage, which were reviewed by humans before being put into the catalog.

    4. Frontend development
    An Angular application was created to consolidate Dashboards, Reports, Data Sets, and Connections into a single architecture. Highcharts was used to visualize revenue and performance data. UX consistency testing was done in all four modules before release.

    5. QA and test automation
    Two QA teams carried out testing parallel to stabilization and module development and not as a post-development phase. The process of GenAI helped broaden the scope of testing without increasing the number of heads in the team, since the focus was primarily on revenue-related features due to their direct connection to publisher trust.

  • The selected stack enabled automated metadata management, scalable BI integration, and AI development:

    Data engineering: Python, Apache Airflow, PostgreSQL

    Metadata management: OpenMetadata, Atlan, AI agent for metadata generation

    Internal tooling: Custom CLI (Python), GitLab CI/CD, Mise

    Frontend: TypeScript, Angular, Highcharts

    QA and automation: Test automation frameworks, GenAI-assisted testing

    Infrastructure: Kubernetes, Docker, GitLab CI/CD

Value Delivered

The publisher’s data team now has a platform that scales without manual rework at every step.

  1. Stable internal tooling layer

    Dependency conflicts that previously blocked the pipeline are now resolved.

  2. Systematic metadata management

    Lineage is now documented automatically and systematically for every data source.

  3. Reduced engineering workload

    The AI agent generates Atlan descriptions automatically, with human review replacing manual write-up.

  4. Unified analytics interface

    One frontend across Dashboards, Reports, Data Sets, and Connections replaces four separate workflows.

Project Results

The rebuilt platform delivered measurable gains across automation, tooling, and reliability.

  • Lineage between data sources and transformations imports into OpenMetadata automatically, removing manual documentation.

  • The AI agent generates Atlan import files, with source, table, column, and lineage descriptions created without a fully manual process.

  • The internal CLI tool distributes across projects through GitLab CI/CD and Mise. Updates reach teams automatically.

  • One unified Angular architecture replaced four previously separate frontend module structures.

  • Dependency conflicts are resolved, and the Python, Airflow, and PostgreSQL pipeline runs predictably.

  • Zero
    Manual Lineage Overhead

  • 100%
    CI/CD Tool Distribution

  • Stable
    Data Pipeline Delivery

Banner

Have a project? Let’s get to  work!

Please enter your name
Please enter valid email address
Please enter from 25 to 500 characters
Required field

Thank you for sharing your needs with us!

We will contact you within 24 hours to discuss your project in more detail.

We couldn’t process your request

“Something went wrong. Please try again later.”

Get the AI guide that helps you make smarter business decisions.

Plus, join 3,000+ readers receiving practical tech and business tips — no spam, just results.

BONUS FREE Guide

"Do You Really Need AI for Your Business?"

Please enter a valid name
Please enter a valid email
Required field

Thank you for Subscribing!

We couldn't process your request

Something went wrong. Please try again later.