Description
With 1,000+ intelligence professionals serving over 1,900 clients worldwide, Recorded Future is the world’s most advanced, and largest, intelligence company! We're looking for a hands-on Product Manager to own the lifecycle of Recorded Future's AI agents and MCP tools — the intelligent workflows and tool integrations that apply large language models to real threat intelligence problems for our customers. This is not a role focused on training models or building LLMs. Instead, you'll build, evaluate, and ship agents that orchestrate existing models, and shape the MCP tools those agents and our customers rely on to access Recorded Future intelligence. You'll be responsible for the full arc of an agent — from initial build, through evaluation and iteration, to customer testing and deployment — as well as the quality and coverage of our MCP tool surface. You'll write and maintain evals, refine tool descriptions, identify gaps in tool coverage, analyze how tools are actually used, and make the call on what ships. This role suits someone who is technically fluent, comfortable getting their hands dirty in GitHub and prompt engineering, and grounded in strong product judgment about what's worth building. This isa mid-level role for a builder who moves fast, tests rigorously, and cares more about whether an agent or tool solves the customer's problem than whether it demos well. What You'll Do: Own the end-to-end lifecycle of AI agents — design, build, evaluation, customer validation, and deployment. Build agents that orchestrate LLMs and tools against real intelligence use cases, selecting the right model for each task based on capability, latency, and cost tradeoffs. Own and refine the MCP tool surface — writing clear, effective tool descriptions, identifying gaps in coverage, and improving how tools expose Recorded Future intelligence to agents and customers. Analyze MCP tool usage patterns to understand what customers and agents actually invoke, where tools fail or underperform, and where new tools are needed. Write, maintain, and expand evaluation suites to measure agent and tool quality, catch regressions, and guide iteration; update agents and tools as models, data, and customer needs evolve. Test agents and tools directly with customers, gathering feedback and confirming that outputs meet their workflows and expectations before and after launch. Work hands-on in the codebase (GitHub) alongside engineers — reviewing changes, prototyping, and contributing to agent logic and tool definitions where appropriate. Maintain a working understanding of the LLM landscape, tracking the strengths, weaknesses, and cost profiles of available models to make informed build decisions.