About this role
Role Purpose
- The Senior Platform/Solution Architect – Capacity Planning, Performance & Cloud Infrastructure Architecture is responsible for translating business demand, application traffic and workload characteristics into quantifiable infrastructure requirements across microservices, Kubernetes/OpenShift, cloud and on-premises environments.
- The role provides the technical capability to determine CPU, memory, pod/replica, node, cluster, database, storage and network requirements while ensuring performance, scalability, resilience, availability and cost efficiency.
Core Responsibilities
- Lead capacity planning and infrastructure dimensioning for applications, platforms and microservices-based services.
- Translate business growth, transaction volumes and traffic forecasts into infrastructure capacity requirements.
- Develop quantitative workload models covering normal, peak, burst and exceptional traffic conditions.
- Determine appropriate CPU, memory, pod/replica, node and cluster requirements for application services.
- Develop capacity forecasts and infrastructure roadmaps covering short-, medium- and long-term demand.
- Ensure capacity plans support availability, resilience, disaster recovery and business continuity requirements.
- Provide architecture and capacity recommendations for both cloud and on-premises environments.
- Review existing environments to identify over-provisioning, under-provisioning, bottlenecks and capacity risks.
Microservices Capacity Planning & Dimensioning:
- Assess resource consumption and performance characteristics of individual microservices.
- Determine minimum, normal and maximum pod/replica requirements based on workload and service-level objectives.
- Define CPU and memory requests and limits for containers.
- Assess horizontal and vertical scaling requirements and define appropriate scaling policies.
- Determine node density, resource utilisation and cluster capacity requirements.
- Account for service-to-service communication, platform overhead and infrastructure reserve capacity.
- Establish repeatable sizing methodologies for new applications and services.
- Validate sizing assumptions through performance and capacity testing.
Capacity Planning Parameters & Metrics:
- Define and maintain standard parameters for application and infrastructure capacity planning.
- Analyse requests per second (RPS), transactions per second (TPS), concurrent users, sessions and transaction volumes.
- Analyse average, peak and burst traffic and associated growth patterns.
- Assess CPU utilisation, CPU consumption per transaction, memory utilisation, memory peaks and application heap requirements.
- Assess pod counts, replica requirements, scaling thresholds and scaling response times.
- Determine node CPU, node memory and allocatable cluster capacity.
- Assess database TPS, connections, CPU, memory, IOPS and throughput.
- Assess storage capacity, IOPS, throughput and growth.
- Assess network bandwidth, latency and packet rates.
- Factor in high availability, N+1/N+2 resilience, disaster recovery, growth headroom and operational reserve.
Performance Engineering:
- Lead performance engineering and capacity validation for critical applications and platforms.
- Define and oversee load, stress, endurance, spike, scalability and capacity testing.
- Analyse throughput, response time, latency, concurrency and resource utilisation.
- Identify application, platform, database, storage and network bottlenecks.
- Establish performance baselines and capacity thresholds.
- Use performance test results to validate CPU, memory, pod, node and cluster sizing.
- Work with engineering teams to optimise resource consumption and application performance.
Observability & Data-Driven Capacity Planning:
- Use production telemetry and historical data to drive capacity decisions and improvements.
- Implement dashboards and reports to monitor resource utilisation, performance and capacity trends.
- Collaborate with SRE/Platform teams to ensure observability and proactive capacity management.
Required Qualifications (as implied by responsibilities):
- Strong experience in capacity planning, performance engineering, and cloud/on-premises infrastructure design.
- Experience with microservices architectures, Kubernetes/OpenShift, and hybrid cloud environments.
- Ability to translate business demand into actionable infrastructure requirements.
- Familiarity with workload modelling, SLOs/SLIs, and scaling policies.
- Knowledge of telemetry, monitoring, and observability practices.
Industry
- Software / Technology