People Finders Logo

People Finders

Director, Infrastructure & IT Operations

Posted 5 Hours Ago
Remote
Hiring Remotely in USA
170K-180K Annually
Expert/Leader
Remote
Hiring Remotely in USA
170K-180K Annually
Expert/Leader
Leads cloud infrastructure, SRE, production operations, corporate IT, helpdesk, and bot mitigation. Owns AWS architecture, reliability, security, observability, incident response, disaster recovery, identity management, vendor relationships, budgets, and cloud costs. Builds operational standards, SLOs, automation, infrastructure roadmaps, and high-performing teams while improving customer-facing platform availability, traffic protection, employee technology support, and cybersecurity readiness.
The summary above was generated by AI

PeopleFinders.com, the premier online service for consumers to locate, contact and verify people and businesses. Over the past couple of decades the Company has quietly become one of the largest owners of public records data in the country, distributing its products over a vast network of websites.

PeopleFinders | Remote |

Salary - $170k - $180k

About the Role

We are seeking a strategic and hands-on Director, Infrastructure & IT Operations to lead the technology capabilities that keep our digital products, cloud platforms, and employees secure, reliable, and productive.

This leader will oversee cloud infrastructure, Site Reliability Engineering, production operations, bot mitigation, corporate IT, and helpdesk support, while providing architectural leadership across cloud platforms, networking, identity, observability, and operational tooling.

The ideal candidate combines strong technical judgment with disciplined operational leadership: building high-performing teams, setting clear service expectations, improving reliability, managing vendors and costs, and reducing operational risk across the company's portfolio of digital properties.

This is not a traditional corporate IT position — it spans both internal employee technology and the production infrastructure supporting high-traffic, customer-facing applications.

Key Responsibilities

Infrastructure and Technology Leadership

Lead the teams responsible for cloud infrastructure, SRE, production operations, corporate IT, helpdesk, and bot mitigation. Establish ownership, operational standards, SLOs, escalation procedures, and performance metrics for each function. Develop managers, technical leads, and IT staff through coaching and career development. Build quarterly and annual infrastructure roadmaps aligned with business priorities, growth, security, reliability, and cost. Partner with Engineering, Product, Data, Marketing, Finance, HR, Legal, and executives, providing clear recommendations on risks, architecture, and tradeoffs. Build a culture of accountability, documentation, automation, and continuous improvement.

Cloud Infrastructure and Architecture

Own the reliability, scalability, security, performance, and cost effectiveness of the company's cloud infrastructure, primarily AWS. Provide architectural leadership across cloud services, applications, APIs, networking, identity, databases, storage, observability, and integrations. Establish architecture standards, reference patterns, governance, and technical review processes. Partner with engineering to improve deployment safety, performance, and readiness. Drive automation, configuration management, and infrastructure as code. Lead capacity planning for business growth and demand spikes. Eliminate single points of failure, undocumented systems, manual processes, and reliance on individual employees or vendors. Review new systems for architectural fit, supportability, security, and total cost of ownership. Ensure production and nonproduction environments are properly separated, secured, and monitored.

Site Reliability and Production Operations

Establish and maintain SLIs, SLOs, availability targets, and error-budget practices for critical systems. Improve observability through centralized logging, metrics, distributed tracing, APM, and actionable alerting. Lead major incident management, including coordination, executive communication, root-cause analysis, and corrective-action tracking. Develop and test backup, disaster recovery, and business continuity procedures. Improve mean time to detect, respond, and recover from incidents. Establish on-call practices balancing coverage with team sustainability. Partner with development teams to improve resilience and production support. Use incident data to prioritize long-term reliability work. Ensure systems have current runbooks, architecture diagrams, and documentation.

Bot Mitigation and Traffic Protection

Own the strategy for bot mitigation, scraping protection, automated abuse prevention, and traffic-quality management. Protect company properties from scraping, credential attacks, fraud, automated abuse, and infrastructure exhaustion. Lead configuration of technologies such as Cloudflare, DataDome, WAFs, rate limiting, CAPTCHA alternatives, device fingerprinting, and application-level defenses. Partner with Engineering, Product, Marketing, SEO, and Revenue to distinguish malicious automation from legitimate customers, search engines, partners, and approved AI crawlers. Establish monitoring and response procedures for scraping and abnormal traffic. Measure effectiveness through traffic reduction, false-positive rates, platform stability, and infrastructure savings. Evaluate vendors for measurable value, and stay current on scraping techniques, browser automation, residential proxies, and AI crawlers.

Corporate IT and Helpdesk

Oversee internal IT services and helpdesk support for employees. Establish service-level expectations for response time, resolution time, satisfaction, and backlog reduction. Own the employee technology lifecycle: onboarding, offboarding, equipment provisioning, access management, and asset recovery. Manage laptops, endpoints, collaboration platforms, productivity software, and telecom services. Standardize endpoint configuration, encryption, patching, and endpoint detection. Improve the support experience through clear intake, self-service resources, and automation. Maintain accurate inventories of devices, licenses, and assets. Partner with HR so new employees get equipment and access on time and departing employees have access promptly removed. Identify and address the root causes of recurring support issues.

Identity, Access, and Operational Security

Establish consistent identity and access-management practices across corporate and production systems. Implement role-based access, least-privilege principles, MFA, periodic access reviews, and timely access removal. Maintain secure processes for privileged, administrative, and service accounts. Partner with security and legal to identify and reduce technology risks. Support vulnerability remediation, security assessments, audits, and vendor reviews. Ensure cloud environments, endpoints, applications, and third-party services meet company security standards, and maintain readiness for cybersecurity incidents and applicable privacy requirements.

Vendor, Budget, and Cost Management

Manage infrastructure, security, IT, telecom, monitoring, and support vendors, including evaluations, contract reviews, renewals, licensing, and SLAs. Develop and manage infrastructure and IT budgets. Improve cloud cost visibility through tagging, allocation, and forecasting. Identify opportunities to eliminate unused services, consolidate tools, and renegotiate contracts. Communicate budget performance and investment recommendations to leadership. Ensure critical vendors have appropriate security controls and documented ownership, and reduce dependency on vendors through internal knowledge.

Leadership Expectations

Operate as both a strategic technology leader and an effective technical decision-maker. Remain calm, organized, and decisive during outages, security events, and high-pressure situations. Create accountability without unnecessary bureaucracy. Communicate technical risks and recommendations clearly to technical and nontechnical audiences. Build productive partnerships across Engineering, Product, Data, Marketing, Security, and business operations. Make decisions based on reliability, risk, customer impact, cost, and measurable outcomes. Develop leaders who can independently manage their functions while maintaining consistent standards, and encourage teams to solve root causes rather than rely on repeated manual intervention.

Required Qualifications

10+ years in cloud infrastructure, SRE, platform engineering, DevOps, IT operations, or related disciplines. 5+ years leading engineering or technology teams, including managers or technical leads. Demonstrated experience operating high-traffic, customer-facing platforms in AWS or a comparable cloud. Strong understanding of cloud architecture, networking, identity, security, observability, databases, APIs, storage, and distributed systems. Experience managing production availability, incident response, root-cause analysis, disaster recovery, and business continuity. Experience with infrastructure automation, configuration management, CI/CD, and infrastructure as code. Experience managing corporate IT, employee support, endpoint management, licensing, and access provisioning. Experience establishing operational metrics, SLOs, and support standards. Demonstrated ability to manage vendors, contracts, budgets, and cloud costs. Strong written and verbal communication skills, with the ability to explain technical issues and tradeoffs to executives.

Preferred Qualifications

Experience with Cloudflare, DataDome, or comparable bot-management and web application protection platforms, and leading bot mitigation or traffic-quality programs. Experience with AWS services, containerized environments, infrastructure as code, centralized logging, and APM (e.g., AppDynamics). Experience supporting large-scale consumer subscription, public-record, data-as-a-service, advertising, or ecommerce platforms. Experience modernizing legacy infrastructure, transitioning systems from external vendors, and consolidating cloud accounts or IT services. Familiarity with privacy requirements and secure handling of sensitive consumer data. Experience building or maturing SRE, DevOps, or platform engineering functions.

Measures of Success

Success will be measured through: improved availability and stability of customer-facing platforms; reduced impact from bots, scraping, and abuse; faster incident detection and recovery; fewer recurring incidents; improved observability and reduced alert noise; tested backup and disaster recovery capabilities; lower cloud waste and better cost visibility; consistent helpdesk performance; reliable onboarding, offboarding, and access management; stronger endpoint security and asset management; reduced reliance on undocumented systems and vendors; stronger architecture governance; and improved team accountability and development

Benefits
401k
401k match
Medical/Dental/Vision/ Life Insurance


Similar Jobs

34 Minutes Ago
Remote
United States
120K-140K Annually
Mid level
120K-140K Annually
Mid level
Computer Vision • Digital Media • Kids + Family • Mobile • Software • Sports
Operate application security across the SDLC, including secure design, code reviews, threat modeling, API security, DevSecOps tooling, CI/CD security gates, cloud and container security, vulnerability management, and risk communication. Partner with engineering teams on secure coding standards, AWS, Kubernetes, Terraform, mobile ecosystems, and responsible AI integration. Participate in an on-call rotation and develop scalable security platforms, paved roads, documentation, and developer training.
Top Skills: Ai/MlAndroidAWSBugcrowdCdnCi/CdGithub ActionsGithub Advanced Security (Ghas)GraphQLInfrastructure As Code (Iac)iOSKotlinKubernetesNowsecureOwasp Top 10RestSwiftTerraformTypescriptWeb Application Firewall (Waf)Wiz
34 Minutes Ago
Remote or Hybrid
Senior level
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Leads presales solution consulting for ServiceNow Integrated Risk Management and GRC solutions. Translates customer processes and pain points into configured solution environments, develops sales campaigns and demonstrations, advises on risk and compliance programs, collaborates with product and development teams, and enables other consultants. Requires strong GRC expertise, enterprise experience, communication skills, and up to 25% travel.
Top Skills: ArcherGovernance Risk And Compliance (Grc)Ibm OpenpagesMetricstreamOnetrustServicenowServicenow Integrated Risk Management
51 Minutes Ago
Remote
USA
Senior level
Senior level
Artificial Intelligence • Machine Learning • Software • Defense
Design, build, and improve scalable backend systems that collect hard-to-access OSINT; build pipelines and ETL for messy data; collaborate with product and mission teams; mentor engineers and ensure reliable data delivery for DoD workflows.
Top Skills: AWSDatabasesDistributed SystemsETLPythonTemporalWeb Scraping

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account