Postgre SQL DBA
Job Description
You will serve as a primary subject matter expert for the enterprise database infrastructure, responsible for the end-to-end administration, performance, and lifecycle management of business-critical PostgreSQL environments. While this role encompasses comprehensive database management duties, a core specialization involves owning the design, implementation, and operation of resilient High Availability (HA) and Disaster Recovery (DR) architectures across shared-storage and streaming-replication models.
Primary Responsibilities
Database Administration & Engineering
* Manage the complete database lifecycle, including standard installations, version upgrades, patching, and secure configuration management.
* Optimize database performance through advanced query tuning, index optimization, and the implementation of declarative table partitioning strategies for large-scale datasets.
* Develop and maintain robust backup and recovery strategies to meet strict Recovery Point Objective (RPO) and Recovery Time Objective (RTO) SLAs.
* Collaborate closely with infrastructure teams on Red Hat Enterprise Linux (RHEL) system administration, focusing on storage provisioning, Logical Volume Manager (LVM) expansion, and general OS-level troubleshooting to ensure optimal database performance.
High Availability & DR Architecture (Specialization)
* Design and implement Patroni and etcd streaming-replication HA architectures for both single-datacenter (2-node plus witness) and cross-datacenter (multi-site plus arbiter) deployment models.
* Design and implement shared-storage Active-Passive clusters utilizing Pacemaker, Corosync, qdevice, and STONITH fencing mechanisms.
* Engineer robust quorum, witness, and arbiter layers customized for each architectural model, explicitly managing per-node versus per-site voting, majority logic, and split-brain prevention.
* Determine and justify the optimal replication mode (synchronous, asynchronous, or sync-with-degrade) and connection-routing strategy (VIP versus multi-host string) based on workload requirements.
* Independently author comprehensive High Level Design (HLD) documentation and successfully secure architecture sign-off from key stakeholders.
Automation & Operational Resilience
* Take full ownership of installer and deployment scripts for HA environments, managing cluster networking, shared storage provisioning, and automated database node setup.
* Apply infrastructure changes by following established design patterns: isolating new logic, documenting every modification, and preserving backward compatibility with existing automation frameworks.
* Configure and monitor critical HA telemetry signals, including replication lag, cluster health, fencing device reachability, and automated failover events.
* Execute rigorous verification tests and failover drills following every system change, investigate behavior post-failover, and remediate underlying issues to restore full architectural redundancy.
Requirements
Function: Information Technology
Experience Level: Mid-Senior Level