As a Senior Software Developer – AI Data Engineer, you will play a key role in defining, designing, and implementing the data architecture that powers Caseware’s AI platform. You will lead architecting and building scalable data ingestion, processing, storage, and retrieval systems across structured, unstructured, and multi-modal data sources, while establishing data models, governance standards, and platform capabilities that support vector search and graph-based knowledge representations. Working across AI, platform, and product teams, you will define and evolve reference architectures, data standards, guardrails, and best practices, while implementing data quality, lineage, traceability, metadata management, freshness monitoring, and alerting capabilities to ensure trusted, reliable, and audit-ready data products that enable AI innovation at scale across Caseware Cloud.
In this role, you will take ownership of complex, production-grade AI data workflows end-to-end, influence architectural direction through technical leadership and proof-of-concepts, and help ensure our Data platform is scalable, reliable, and measurable. You will collaborate closely with AI, platform, DevOps, and product teams to turn emerging AI patterns into durable platform capabilities that directly impact customers across Caseware’s cloud ecosystem.
Caseware is evolving Caseware Cloud to deliver intelligent, data-driven experiences—powering analytics, automation, and AI/agentic capabilities on top of a modern data platform.
This role is for someone who can bridge transactional backend systems and data-intensive distributed workflows. You’ll work on systems that combine:
APIs and domain services (microservices, relational modeling, service boundaries)
Asynchronous workflows (messaging, retries, idempotency, replay safety)
Distributed/batch data processing (Spark-based processing and lake patterns)
Cloud platform primitives (AWS orchestration and managed services)
AI-ready retrieval workflows (embedding + vector retrieval pipelines)
Improved reliability and operability of ingestion + async workflows (clearer idempotency/replay patterns, fewer recurring incidents).
Cleaner boundaries between orchestration/control-plane concerns and data-processing execution concerns.
Better observability across APIs, queues, workflows, and distributed jobs.
Clearer data contracts and more predictable schema evolution practices.
Tangible improvements in developer experience (local run, testing, reduced “environment-only” hacks).