HomeSearchData Engineer Jobs › Senior Data Engineer - MIDAS Data Platform, Digital Bank, Tokyo

Senior Data Engineer - MIDAS Data Platform, Digital Bank, Tokyo

Money Forward

Minato-ku, Tokyo, JP · senior
More Data Engineer jobs: Data Engineer jobsData Engineer salary

Job at a glance

Minato-ku, Tokyo, JP
Location
Senior
Seniority
Money Forward
Employer
Data Engineer jobs
Category

As a Senior Data Engineer in the MIDAS (Management Integration & Data Analytics System) Data Platform Team, you will build from scratch and maintain the central data hub connecting most systems found inside one of Japan’s more innovative digital banks. You will work with modern cloud-based data technologies to ingest data from various banking systems, apply complex business logic on it, then serve it to downstream systems for enterprise management, regulatory reporting, risk management and many other applications.

Thanks to the high expectations towards the banking domain, you will have the opportunity to work on complex data engineering challenges including data quality, reconciliation across multiple systems, time-critical data processing, and complete traceability. This is a senior individual contributor role where you will design and implement complex data pipelines, mentor mid-level engineers, and participate in architectural decisions for the platform.

※ This position involves employment with Money Forward, Inc., and a secondment to the new company (SMBC Money Forward Bank Preparatory Corporation). The evaluation system and employee benefits will follow the policies of Money Forward, Inc. Who we are? We are a startup team partnering with Sumitomo Mitsui Financial Group and Sumitomo Mitsui Banking Corporation to establish a new digital bank. Our mission is to build embedded financial products from the ground up, with a strong focus on supporting small and medium-sized businesses (SMBs).

Technology Stack and Tools Used Cloud Infrastructure AWS (primary cloud platform in Tokyo region) S3 for data lake storage with VPC networking for secure connectivity AWS IAM for security and access management Data Lakehouse Architecture (Databricks) Modern lakehouse architecture using Delta Lake for ACID transactions, time-travel, and schema evolution Columnar storage formats (Parquet) optimized for analytics Bronze/Silver/Gold medallion architecture for progressive data refinement Partition strategies and Z-ordering for query performance Unity Catalog for centralized governance and metadata management Orchestration & Processing (Databricks) Databricks Workflows for managed workflow orchestration Distributed data processing with Apache Spark on Databricks clusters Serverless compute and auto-scaling clusters for cost optimization Streaming and batch ingestion patterns with Databricks AutoLoader Data Transformation (Databricks) Delta Live Tables for declarative ETL pipelines with built-in data quality SQL and Python for data transformations in Databricks notebooks Incremental materialization strategies for efficiency Query & Analytics (Databricks) Databricks SQL for high-performance analytics queries Serverless and auto-scaling SQL warehouses for variable workloads Query result caching and optimization REST APIs for data serving to downstream consumers Direct Delta Lake access for advanced consumers (Athena, Redshift Spectrum, etc.) Data Quality & Governance (Databricks) Automated data quality with Delta Live Tables expectations Cross-system reconciliation and validation logic Fine-grained access control with column/row-level security using Unity Catalog Automated data lineage tracking for regulatory compliance Audit logging and 10-year data retention policies Business Intelligence (Databricks) Databricks SQL Dashboards for internal analytics and monitoring Integration with enterprise BI tools (Tableau, PowerBI, Looker) via Databricks SQL endpoints Development & DevOps Languages: SQL (primary), Python Platform: Databricks (notebooks, workflows, SQL) Version Control: GitHub CI/CD: GitHub Actions Infrastructure as Code: Terraform Monitoring: Databricks monitoring, AWS CloudWatch integration AI-Assisted Development: Claude Code, GitHub Copilot, ChatGPT Responsibilities Design and implement data pipelines to ingest data from multiple source systems (CRM, CBS, CLM, LOS) using REST APIs or database connections Build and maintain Bronze/Silver/G

Search all live jobs — free, no account →