Description

Job Title: Production Support Specialist

Location: Mississauga, ON – Hybrid, onsite 2–3 days/week

Contract: 6 months CTP, with potential extension/conversion

Role Summary

Seeking a Senior Production Support / L3 Engineer to own production stability, data integrity, and DevOps operations for high-criticality Enterprise Risk Management (ERM) platform.

This is an active L3 application/platform ownership role, not a passive support position. The engineer will bridge Production Support, Development, DevOps, Data Engineering, Infrastructure, and Business teams while supporting 17+ applications/pillars.

Key Responsibilities

  • Lead L3 production support, triage, root-cause analysis, problem management, and major incident management across distributed applications and data pipelines.
  • Troubleshoot complex production issues involving Kafka, Oracle, Java/Spring Boot, REST APIs, microservices, data ingestion, and infrastructure.
  • Manage ServiceNow Incident, Problem, Change, PRB, PTASK, and SWAT processes, including RCA and post-incident documentation.
  • Provide L3 support for the Enterprise Risk Data Layer (ERDL) and Data Lake, including pipeline health, data quality, refresh status, reconciliation, and downstream reporting.
  • Perform data reconciliation and data-quality checks across ERDL, KRI APIs, dashboards, and reporting layers.
  • Troubleshoot Kafka producer/consumer issues, consumer lag, schema mismatches, null/field-level issues, and data contract failures.
  • Support OpenShift/Kubernetes, Harness, CI/CD, GitHub/Bitbucket, Lightspeed, and AppDynamics environments.
  • Monitor platform health using AppDynamics, Splunk/ELK/Kibana, Kafka dashboards, Tableau, and Superset.
  • Identify opportunities to automate operational monitoring, alerting, data-quality checks, and ServiceNow incident creation.

Required Qualifications

  • Technology experience, with strong L3 Production Support / Production Engineering experience.
  • Strong ServiceNow ITSM experience, including Incident, Problem, and Change management.
  • Senior-level incident, problem, escalation, and root-cause management experience.
  • Experience with DevOps, CI/CD, deployment, and production release management.
  • Hands-on OpenShift/Kubernetes experience, including pod management and health checks.
  • Experience troubleshooting distributed microservices and data architectures.
  • Working knowledge of Oracle, SQL/read-level queries, materialized views, batch jobs, and reconciliation views.
  • Ability to troubleshoot Java/Spring Boot applications and REST APIs at the application-support level.

Technical Environment

  • Data & Integration: Kafka/JMS, Oracle, ERDL/Data Lake, Tableau, Superset, Elasticsearch, MongoDB/Couchbase
  • Platform & DevOps: OpenShift, Kubernetes, Harness, Lightspeed, GitHub/Bitbucket, CI/CD
  • Application: Java, Spring Boot, REST APIs, Angular/React
  • Monitoring: AppDynamics, Splunk, ELK/Kibana
  • DevSecOps: SonarQube, Snyk, Checkmarx
  • Security: CyberArk, HashiCorp Vault, SSL/TLS, CVM/CAMP, access/entitlement management