- Does this position require a security clearance? No
- Years 3 to 5+ years
- Applicants Less than 10 applicants
- Additional Info Visa / work permit sponsorship is not available for this position
- Applicants are required to read, write, and speak the following languages English
Job Description
Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers’ implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.
Responsibilities
- Implements and contributes to the development for components of distributed systems that support horizontal and vertical scaling including leveraging distributed state management tools.
- Optimizes code and/or systems for large‑scale data processing in large‑scale systems.
- Implements scalability requirements for assigned components and reviews implementation of team members.
- Leverages components of data plane platforms to handle large‑scale data retrieval, storage, and processing.
- Implements performance and load testing.
- Collaborates with team to build fault‑tolerant components capable of withstanding in‑service updates by implementing redundancy, replication, and automatic failover mechanisms.
- Applies recovery‑oriented computing principles to design components that effectively handle service disruptions.
- Implements retry mechanisms, circuit breakers, and timeouts to help handle network unreliability.
- Implements tests and alarm configurations to proactively detect and address issues/failures.
- Supports efforts to recover from failures by drafting and executing runbooks and operational procedures.
- Builds and customizes dashboards, telemetry systems, and alerting mechanisms to monitor component health.
- Designs and implements functional requirements and testing for assigned features within an existing system.
- Implements test scenarios (e.g., fault‑injection, brown‑out) to evaluate system correctness.
- Implements standard data replication and synchronization techniques to maintain data integrity and availability.
- Diagnoses, debugs, and resolves issues in system components to support ongoing operation.
- Implements basic strategies to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
- Designs and implements automation scripts and tooling used to troubleshoot operational issues.
- Participates in operational support rotations, assisting in incident responses and root cause investigations.
- Applies advanced security measures to protect data and applications in multi‑tenant environments, including encryption and access controls.
- Implements remediation plans to continuously improve security.
- Collaborates with the team to ensure cloud infrastructure complies with relevant industry standards and regulations and that documentation is up‑to‑date.
- Maintains automation scripts and tools (e.g., Infrastructure as Code (IaC)) for managing cloud infrastructure.
- Adheres to change management plans for patching, updating, and rolling back applications.
- Tracks timelines with minimal supervision, ensuring work is completed in a timely manner and is in alignment with project requirements.
- Prioritizes and adjusts work as resources or timelines change, with some guidance.
- Collaborates across teams to align on expectations and achieve shared objectives, builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs, and actively listens to diverse perspectives.
- Independently identifies and addresses standard and non‑standard issues in accordance with standard practices, escalating more complex issues as appropriate, analyzing data and/or information from multiple sources to troubleshoot standard and non‑standard errors, and contributing to knowledge sharing and best practices.
- Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools, staying current with industry trends and best practices, seeking out and leveraging feedback and training to improve skills, and contributing to a culture of continuous learning and knowledge sharing with team members.
- Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team, and seeks input from team members on alternative approaches and methods for improving work.
Qualifications
Qualifications are not listed.
We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.
#J-18808-Ljbffr…
