Back to jobs
New

Senior Agentic AI Systems Engineer – Runtime Frameworks & Resource Intelligence

Houston, TX

About HP

At HP, you’ll have a chance to create tools, technology, and solutions that reshape the way the world works in the future. If you’re looking to join a company that allows you to connect with a network of professionals eager to support you in doing your best work, we want to talk to you. Our legendary culture guides every employee toward success—fostering collaboration and driving innovation.

Operating in over 170 countries, HP is always creating new services, products, and capabilities giving you more opportunities to advance your career. Here, innovation is the key to professional development and career mobility.

When you join HP, you’re joining a company that believes every voice matters and that we all deserve a seat at the table. From the boardroom to factory floor, we create a culture where everyone is respected and where people can be themselves. You will be part of a global laboratory where different perspectives and experiences will help you solve problems in new ways. This is where you can build a long and wide-ranging career.

 

The Commercial Systems Software Engineering organization is developing next-generation AI software platforms that enable intelligent, secure, and scalable agentic experiences across client and edge computing environments.

We are seeking an experienced Senior Agentic AI Systems Engineer to analyze, tailor, and integrate agentic AI frameworks for environments where multiple agents operate concurrently under changing system conditions. The role will focus on the runtime frameworks, resource-management policies, system services, and platform interfaces required to deliver responsive and reliable agentic experiences across diverse PC and edge platforms.

The successful candidate will combine knowledge of agentic AI frameworks with strong systems software, performance engineering, memory management, and heterogeneous computing experience. The individual will evaluate how agentic workloads interact with CPU, GPU, NPU, memory, storage, network, battery, power, and thermal resources, and develop adaptive capabilities that enable effective execution within available system constraints.

Responsibilities

  • Analyze agentic AI frameworks and runtime architectures for deployment across PC and edge computing environments.

  • Tailor and integrate agentic frameworks to operate effectively under changing resource and system conditions.

  • Evaluate how multiple concurrent agents, AI models, and applications use and compete for compute, memory, storage, networking, power, and thermal resources.

  • Develop policies for workload admission, prioritization, scheduling, orchestration, throttling, suspension, recovery, and termination.

  • Create capabilities that maintain application and system responsiveness while multiple agentic workloads execute concurrently.

  • Develop graceful-degradation strategies for memory pressure, compute contention, low battery, thermal limits, intermittent connectivity, and other constrained conditions.

  • Analyze the system impact of agent planning, tool execution, retrieval, model inference, context management, state persistence, and inter-agent communication.

  • Define approaches for managing agent context, memory, state, shared resources, and long-running sessions.

  • Develop abstractions that support different AI models, runtimes, hardware configurations, and local, edge, or cloud-connected execution environments.

  • Create workload characterization, benchmarking, stress-testing, and validation methods for multi-agent scenarios.

  • Instrument agentic runtimes to collect telemetry related to performance, resource consumption, responsiveness, reliability, and execution outcomes.

  • Diagnose complex issues spanning agentic frameworks, AI runtimes, operating systems, applications, services, and hardware resources.

  • Develop reusable runtime services, software components, APIs, libraries, tools, and platform integration guidance.

  • Collaborate with architects, AI engineers, systems engineers, application teams, quality teams, and technology partners to deliver production-ready capabilities.

  • Apply secure development, automated testing, observability, CI/CD, and operational best practices throughout the software lifecycle.

Education & Experience Recommended

  • Four-year or graduate degree in Computer Science, Software Engineering, Computer Engineering, Electrical Engineering, or a related discipline, or equivalent experience.

  • 8+ years of experience in systems software, AI infrastructure, runtime engineering, performance engineering, distributed systems, or a related field.

  • Demonstrated experience evaluating, integrating, or adapting AI frameworks, runtime systems, orchestration platforms, or complex software platforms.

  • Experience developing performance-sensitive software for client, edge, embedded, or other resource-constrained environments.

  • Experience analyzing system-level performance, scalability, reliability, and resource-contention challenges.

  • Proven ability to diagnose issues involving complex interactions across applications, operating systems, AI runtimes, and hardware.

  • Experience delivering software in collaboration with architecture, engineering, product, quality, and partner organizations.

Required Technical Skills

Agentic AI Frameworks and Runtime Systems

  • Strong understanding of agentic AI concepts, including planning, reasoning, memory, retrieval, tool use, state management, workflow orchestration, and multi-agent coordination.

  • Experience with agentic frameworks and orchestration technologies such as Semantic Kernel, AutoGen, LangGraph, LangChain, LlamaIndex, or comparable platforms.

  • Understanding of sequential, parallel, event-driven, hierarchical, and long-running agent execution patterns.

  • Familiarity with tool integration, agent communication, shared services, workflow state, and agent lifecycle management.

  • Understanding of how planning loops, model calls, retrieval operations, tool execution, and communication patterns affect latency and resource utilization.

  • Experience evaluating framework behavior, scalability, reliability, recovery, and runtime efficiency.

  • Understanding of checkpointing, retries, fault recovery, state restoration, and graceful failure within agentic workflows.

  • Familiarity with retrieval-augmented workflows, multi-model pipelines, and local, edge, and cloud-connected AI execution.

Resource and Workload Management

  • Strong understanding of workload scheduling, resource governance, admission control, workload isolation, prioritization, and Quality of Service principles.

  • Experience analyzing concurrent workloads across multiple agents, applications, models, services, and system processes.

  • Ability to define resource budgets and operating policies for CPU, GPU, NPU, memory, storage, network, power, and execution time.

  • Experience designing adaptive systems that respond to changing workload demands, system health, battery state, thermal headroom, and resource availability.

  • Understanding of asynchronous execution, concurrency, event-driven systems, inter-process communication, and distributed coordination.

  • Ability to balance agent responsiveness with foreground activity, system stability, battery life, and overall user experience.

  • Experience developing policies for queuing, throttling, pausing, resuming, migrating, or terminating lower-priority workloads.

  • Ability to establish runtime guardrails that protect critical workloads and maintain consistent system behavior.

Memory, Context, and State Management

  • Strong understanding of memory usage across agent frameworks, AI models, runtime services, applications, and operating systems.

  • Understanding of transformer context windows, token utilization, KV cache behavior, and the memory implications of long-running or concurrent agent sessions.

  • Experience analyzing context growth, retained state, conversation history, retrieval results, tool outputs, and shared resources.

  • Ability to define policies for context retention, summarization, compression, reuse, eviction, persistence, and recovery.

  • Familiarity with short-term, episodic, semantic, and persistent memory patterns used in agentic systems.

  • Experience balancing context continuity and quality against memory consumption, latency, privacy, and storage requirements.

  • Understanding of model residency, dynamic model loading, shared model services, caching, paging, and memory pressure.

  • Ability to design reliable approaches for sharing context and state across agents while maintaining appropriate security and isolation.

Performance, Power, and Thermal Engineering

  • Strong experience with performance profiling, workload characterization, benchmarking, diagnostics, and root-cause analysis.

  • Ability to evaluate agentic workloads using latency, throughput, responsiveness, reliability, resource utilization, and energy-efficiency metrics.

  • Experience analyzing CPU, GPU, NPU, memory, storage, network, battery, power, and thermal telemetry.

  • Understanding of how sustained AI workloads affect power consumption, battery life, thermal conditions, frequency scaling, and overall system responsiveness.

  • Ability to define runtime behavior that adapts to power mode, thermal headroom, memory pressure, foreground activity, and workload priority.

  • Experience evaluating systems under sustained load, concurrent usage, resource contention, changing connectivity, and platform stress.

  • Ability to identify bottlenecks across agent frameworks, AI runtimes, operating systems, applications, and hardware resources.

  • Experience translating performance findings into framework changes, resource policies, and platform recommendations.

AI Runtime and Software Development Skills

  • Experience with AI runtimes and execution frameworks such as ONNX Runtime, OpenVINO, DirectML, TensorRT, or comparable technologies.

  • Understanding of heterogeneous computing across CPUs, GPUs, NPUs, memory, storage, networking, and cloud resources.

  • Familiarity with model inference characteristics, including latency, throughput, memory footprint, context capacity, and hardware compatibility.

  • Experience integrating AI services with operating systems, applications, device services, and cloud-connected platforms.

  • Familiarity with APIs, service architectures, messaging systems, event-driven platforms, containers, and distributed software systems.

  • Advanced proficiency in Python and C++.

  • Strong experience in systems programming, multithreading, concurrency, asynchronous execution, and software integration.

  • Experience developing runtime services, middleware, SDKs, APIs, libraries, or other reusable platform components.

  • Ability to analyze logs, traces, memory profiles, performance counters, crash data, and system telemetry.

  • Experience with Git, code reviews, automated builds, CI/CD pipelines, DevOps practices, automated testing, and release processes.

  • Understanding of software security, privacy, permissions, policy enforcement, dependency management, and operational reliability.

Leadership and Collaboration Skills

  • Strong systems thinking and ability to solve problems across AI, software, operating system, and hardware boundaries.

  • Ability to work effectively with application and agent-development teams to integrate resource-aware runtime capabilities into broader agentic solutions.

  • Strong collaboration skills across architecture, software engineering, AI engineering, performance, quality, security, product, and program teams.

  • Ability to clearly communicate complex system behavior, architectural tradeoffs, performance findings, and technical recommendations.

  • Strong technical writing and presentation skills, including design documentation, architecture reviews, and engineering guidance.

  • Ability to influence technical decisions and drive alignment across teams without direct authority.

  • Ability to mentor engineers on agentic runtime architecture, performance analysis, debugging, and systems engineering practices.

  • Customer-focused approach with an emphasis on responsiveness, reliability, efficiency, scalability, and consistent user experiences.

  • Ability to operate effectively in an evolving technology area with changing requirements and platform capabilities.

Preferred Qualifications

  • Experience deploying agentic AI frameworks on PCs, edge devices, embedded platforms, or other resource-constrained systems.

  • Experience adapting open-source or commercial agentic frameworks for production environments.

  • Experience developing resource-aware runtimes, schedulers, orchestration systems, policy engines, or workload-management services.

  • Familiarity with operating-system scheduling, power management, thermal management, memory management, or resource-governance technologies.

  • Experience with long-running AI workflows, context management, agent memory, retrieval-augmented generation, or shared model services.

  • Familiarity with unified-memory architectures and shared-memory accelerator designs.

  • Experience building telemetry, observability, performance-analysis, or system-health capabilities.

  • Experience collaborating with silicon providers, operating-system vendors, cloud providers, or AI ecosystem partners.

  • Patents, publications, open-source contributions, or demonstrated innovation in AI runtimes, agentic frameworks, resource management, or systems software.

Impact & Scope

This role will help define how multiple AI agents operate effectively across future PC and edge platforms. The engineer will develop the runtime frameworks, system services, resource policies, and observability capabilities required to balance agent execution with performance, memory availability, battery life, thermal limits, reliability, and foreground user experiences.

The resulting capabilities will enable agentic applications to adapt to different device configurations and operating conditions while supporting scalable and reliable execution across diverse customer scenarios.

Complexity

Solves complex problems at the intersection of agentic AI frameworks, runtime systems, memory and context management, operating systems, heterogeneous computing, performance, power, and thermals. The role requires balancing agent responsiveness and capability against constrained and continuously changing system resources while supporting multiple concurrent workloads, applications, models, and platform configurations.

Salary: $154,400 - $227,750 

 

Compensation & Benefits (Full-Time Employees)

The salary range for this role is listed above. Final salary offered is based upon multiple factors including individual job-related qualifications, education, experience, knowledge and skills.

At HP, we offer a competitive and comprehensive benefits package, including:

  • Health insurance
  • Dental insurance
  • Vision insurance
  • Long term/short term disability insurance
  • Employee assistance program
  • Flexible spending account
  • Life insurance
  • Generous time off policies, including; 
    • 4-12 weeks fully paid parental leave based on tenure
    • 11 paid holidays
    • Additional flexible paid vacation and sick leave (US benefits overview)

Why join HP?

When you join HP, you’re investing in your future—and so are we. You have priorities beyond work. That’s why we offer flexibility, support, and benefits that help you shape the life you want. 

Equal Opportunity Employer (EEO) Statement

HP, Inc. provides equal employment opportunity to all employees and prospective employees, without regard to race, color, religion, sex, national origin, ancestry, citizenship, sexual orientation, age, disability, or status as a protected veteran, marital status, familial status, physical or mental disability, medical condition, pregnancy, genetic predisposition or carrier status, uniformed service status, political affiliation or any other characteristic protected by applicable national, federal, state, and local law(s).

Please be assured that you will not be subject to any adverse treatment if you choose to disclose the information requested. This information is provided voluntarily. The information obtained will be kept in strict confidence.

Apply for this job

*

indicates a required field

Phone
Resume/CV

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Select...
Select...