Tools for the hard parts of forward deployed engineering.
Customer-specific integrations, imperfect data, real permissions, and a result someone else can operate.
A forward deployed engineer works beside the customer to turn an actual business workflow into working software. The awkward parts matter: two systems disagree about a record, consent expires halfway through a sync, or a retried write creates a second transaction.
This collection focuses on those problems. A tool earns its place through a specific field use and a meaningful difference from its neighbors. Familiar tools can qualify; popularity alone cannot. General developer-stack essentials are outside the scope.
- Carry context across the engagement 1 tool
- Untangle workflows and records 9 tools
- Build the missing integration 8 tools
- Govern agent tool access 2 tools
- Preserve access across systems 7 tools
- Actuate browsers and live customer sessions 3 tools
- Isolate untrusted agent execution 2 tools
- Review data changes before cutover 9 tools
- Reproduce hostile dependencies 10 tools
- Extract difficult customer inputs 7 tools
- Survive retries and human approvals 6 tools
- Prove the customer task works 9 tools
Start with the failure you need to reproduce or the evidence you need to leave behind. These are alternatives to evaluate, not a stack to install in full.
Keep discovery decisions, evidence, and agent behavior available at cutover and to the next engineer.
FDEOps agent engagement skills
Install task skills and a coordinator that give an AI coding agent an engagement workflow - discover, build, hand off - with a local per-customer memory so decisions and evidence survive between sessions.
Field note: Skills are instructions an agent executes. Review them before pointing them at customer material, and check data-handling rules: the local memory store holds customer context.
Find where the work stalls and which records represent the same real-world thing.
-
PM4Py process mining
Discover process variants and check conformance from customer event logs. Useful when the workflow described in interviews differs from the paths cases actually take.
Field note: Needs usable case identifiers, activities, and timestamps. A diagram cannot repair an incomplete event log. -
Splink probabilistic linkage
Link customer, supplier, or patient records across systems without a shared identifier. Fellegi–Sunter match weights and comparison charts make the matching decision inspectable.
Field note: Choose blocking rules and review borderline matches. A high match probability is not proof that two people are the same. -
dedupe human-labeled matching
Train record matching with examples labeled by the customer's domain expert. Active learning helps turn their judgment about names, addresses, and duplicates into a repeatable matcher.
Field note: Choose this when labeled examples should drive the model; use Splink when you need its probabilistic diagnostics and SQL backends. -
OpenRefine operator-led cleanup
Cluster inconsistent values, reconcile them against a reference service, and retain the cleanup operations. Useful when an operator must approve ambiguous spreadsheet corrections.
Field note: Interactive cleanup comes first. Export the settled rules and test them before turning the session into an unattended import. -
qsv CSV investigation
Inspect large customer CSVs, infer a JSON Schema, calculate statistics, and isolate malformed records through a CLI. Useful before trusting an export or writing ingestion code.
Field note: Memory use depends on the command and options. Cardinality, quantiles, and schema enums do not share streaming commands' constant-memory behavior. -
VisiData jump-host data exploration
Inspect, join, filter, and compare customer extracts in a terminal spreadsheet. Useful on a remote host where a desktop cleanup application cannot run.
Field note: An interactive investigation tool, not an unattended validator. Keep the final transformation rules and test inputs outside the session. -
marimo rerunnable analysis notebooks
Run discovery-week analysis as reactive notebooks stored as plain Python that git-diff cleanly and re-execute deterministically. Useful when the exploration behind a data decision must stay reviewable through handoff.
Field note: Reactivity removes stale-cell surprises but re-executes downstream cells on edit. Pin data snapshots when rerunning against changed customer data would confuse the record. -
Steampipe live API investigation
Query customer SaaS and cloud accounts through SQL plugins without first building a warehouse. Useful for checking cross-system inventory and permission assumptions during discovery.
Field note: Plugin calls still need approved credentials and API budgets. Results are live or cached, not a durable historical replication pipeline. -
Datasette explorable extract handoff
Publish a customer extract or migration dataset as a read-only web UI and JSON API with faceted search. Useful when operations staff need to find and verify records without waiting for an application to be built.
Field note: A publishing layer, not a secured application. Put it behind the customer's access controls; the interface does not make a sensitive extract safe to expose.
Handle customer-specific grants, extraction state, and legacy interfaces instead of assuming a ready-made connector.
-
Nango OAuth lifecyclesource-available
Manage per-customer API connections, token refresh, and authenticated proxy requests. Useful when each customer grants access separately and revoked consent needs a visible reconnect path.
Field note: Free self-hosting covers Auth and Proxy. Broader production functions and syncs have different deployment and plan requirements. -
Airbyte Python CDK custom source connectors
Build the source connector missing from a customer's Airbyte setup. HTTP streams, pagination, partitions, and incremental state give the next engineer a defined connector interface.
Field note: The connector SDK earns the slot, not a new platform installation. Test cursors, deleted records, and resumption after a partial sync. -
dlt schema-aware loading
Write custom extraction and incremental loads while making schema changes explicit through schema contracts. Freeze an agreed schema or define how unexpected rows and columns are handled.
Field note: Schema inference is not business approval. Decide whether drift should stop the load, evolve the schema, or reject data. -
Apache Camel legacy protocol routing
Connect existing enterprise interfaces such as SFTP, JMS, and SOAP with routing, transformations, and error handling. Useful when the integration boundary is more than a JSON API.
Field note: Check the exact component and endpoint semantics. Route retries still need a business-level duplicate-write strategy. -
Svix signed webhook delivery
Deliver customer-facing webhooks with signatures, retries, attempt visibility, and recovery controls. Useful when a customer endpoint goes offline and needs a reliable way to catch up.
Field note: Delivery is at least once. Receivers must verify signatures and deduplicate; a successful HTTP response does not prove the business transaction completed. -
Redpanda Connect stream normalization
Map and route heterogeneous messages through declarative pipelines and Bloblang. Useful when customer queues and webhook payloads need normalization rather than another batch ELT connector.
Field note: The successor to Benthos. Check the chosen components and edition; at-least-once stream delivery and logs do not provide per-record provenance. -
Apache NiFi operator-visible dataflows
Build queued flows with backpressure and searchable record provenance. Useful when customer operators need to inspect where an individual record went and why it stopped.
Field note: An operated service, not a laptop utility. Configure provenance retention and disk budgets; prefer the customer's existing installation when available. -
FastMCP MCP server scaffolding
Build an MCP server over a customer system that has no agent surface, with typed tools, auth middleware, and testing utilities instead of hand-written protocol plumbing. Useful when the integration deliverable is an action interface for agents rather than another extraction pipeline.
Field note: Wrapping a legacy API adds no authorization. Map tools to the customer's permission model and test denied calls, not only the happy path.
Keep credentials and tool policy out of the model when an agent must act in the customer's systems.
-
Arcade authorized MCP tool runtime
Build MCP tools that declare their OAuth requirements so the runtime vaults tokens, runs the user consent flow, and injects credentials server-side; the model never receives the secret. Useful when each customer user must authorize agent actions separately.
Field note: The framework is open source; production vaulting and gateway features are commercial. Different job from FastMCP scaffolding: Arcade is the governed execution and per-user authorization layer. -
ToolHive isolated MCP runtime and gateway
Run MCP servers in containers behind a registry, gateway, and identity integration so customer security can approve which tools agents may call and audit the invocations. Useful when an engagement must expose customer systems as tools under IT review.
Field note: Adds platform surface. Start with only the servers the engagement needs; container isolation does not validate tool arguments or replace authorization.
A valid OAuth token and an indexed document do not establish permission to perform the customer's action.
-
SpiceDB relationship permissions
Model resource sharing and inherited permissions across customer systems. ZedTokens let a subsequent check require permission data at least as fresh as a known relationship change.
Field note: You must synchronize source relationships. Default low-latency reads can be stale; test revocation using the required consistency level. -
Cerbos attribute-based policy
Evaluate versioned policies against principals, resources, and actions. Its query-plan API helps filter a list of customer records according to the same policy used for individual checks.
Field note: Provide trustworthy identity and resource attributes. Cerbos is a policy decision point, not an importer for the customer's document ACLs. -
OpenFGA testable sharing models
Model inherited resource access and keep Check, ListObjects, and ListUsers assertions with fixture tuples. Useful for reviewing the customer's sharing rules before the integration goes live.
Field note: Source tuple synchronization remains your job. Higher-consistency reads bypass caches; they are not SpiceDB's write-token freshness mechanism. -
Cedar embedded policy decisions
Evaluate fine-grained policies in an application and validate them against a declared schema. Useful when customer authorization rules must remain separate from code without adding a policy server.
Field note: The application supplies authenticated principals, entities, and context. Schema validation cannot prove those inputs describe the right customer. -
Pomerium clientless internal-app access
Put identity- and context-based access checks in front of a customer-facing internal web application. Useful when browser users need controlled access without a VPN client.
Field note: The proxy must be reachable. This is application access, not document ACL import or a network that needs no inbound service ports. -
OpenZiti private-service reachability
Connect customer services through identity-based networking with outbound service registration. Useful when a private customer service cannot expose an inbound listening port.
Field note: Requires an operated fabric and a tunneler or embedded SDK. Routers still need reachable endpoints; network access does not grant business-record permission. -
Conftest delivery-config policy checks
Test structured deployment and integration configuration with Rego assertions. Keep customer-specific rules for approved destinations, required settings, and forbidden configuration in the handoff checks.
Field note: Checks declared artifacts before use. It is not runtime authorization and cannot prove that the installed service follows its configuration.
Operate inside the customer's live web apps and give agents a real browser when the workflow has no API.
-
Browserbase hosted agent browsers
Provision isolated cloud Chromium sessions with contexts and proxies so browser agents and demos do not depend on a laptop browser. Useful when a customer UI must be automated at scale or from a server.
Field note: Commercial hosted infrastructure. Confirm data residency, recording retention, and whether customer credentials may enter the session. -
Stagehand self-healing browser agent SDK
Combine deterministic Playwright-style actions with natural-language act, observe, and extract so automations survive customer UI changes better than selectors alone. Useful when a legacy admin UI must be driven without its own API.
Field note: Model-backed steps are non-deterministic and cost calls. Keep critical writes on deterministic locators and record failing pages as fixtures. -
Webfuse
live customer session proxy
Proxy a customer web app into shared Spaces and Sessions so an engineer or agent can co-browse, transfer control, and run guided workflows without a browser extension. Useful when acting inside the customer's SSO-bound SaaS UI with real permissions intact.
Field note: Commercial, and maintained by this list's author. Treat the proxy as a trust boundary: configure masking and audit for sensitive fields, and confirm the customer's security review accepts session capture.
Run model-generated code and computer-use desktops away from the customer's laptop and yours.
-
E2B Firecracker agent sandboxes
Start hardware-isolated Linux microVMs per agent session, with pause, resume, and an optional desktop, for code execution and computer-use inside a customer-approved boundary. Useful when a delivered agent must run code without touching customer or engineer machines.
Field note: Isolation tier and network egress are the decision. A sandbox with open egress can still exfiltrate; configure secrets injection and egress policy deliberately. -
Daytona secure AI code workspaces
Provide disposable or persistent workspaces for AI-generated code with self-host paths suited to customer tenancy. Useful when agents need a real development environment rather than a single execution.
Field note: Default isolation can be container-class; choose VM or hardened tiers for fully untrusted code. Confirm the self-host edition before committing to a customer.
Leave evidence that a schema, migration, or transformation changed what the business intended.
-
Data Contract CLI producer-consumer contracts
Lint and test a shared data contract with schema and quality expectations. Useful when the customer and implementation team need a runnable agreement at their data boundary.
Field note: A passing contract only proves its declared checks. Include ownership and the business assumptions behind required fields. -
Pandera dataframe boundary checks
Validate dataframe columns, types, and custom rules inside ingestion code. Turn the customer's input assumptions into executable checks and identifiable failures.
Field note: Fits code-level validation. Keep representative rejected rows and decide which failures quarantine data versus stop the job. -
Frictionless portable file contracts
Describe and validate tabular drops with Table Schema and Data Package descriptors. Useful when the file, its schema, and a validation report must travel between teams.
Field note: Checks a portable file boundary; it does not replace dataframe rules or warehouse monitoring. -
Recce dbt change review
Compare baseline and candidate dbt environments using lineage, schema, profiles, and query-result diffs. Give business owners concrete before-and-after evidence for a model change.
Field note: dbt-specific. Align source snapshots and time-dependent filters so the review does not mistake changing inputs for changed logic. -
Reladiff cross-database reconciliation
Compare copied tables across database engines with hash-based segment checks, or use a join within one database. Useful for proving that a customer migration retained the expected values.
Field note: Align keys, precision, and snapshots. Engine support varies; two actively changing tables do not provide a consistent comparison. -
OpenLineage run-level provenance
Emit run, job, and dataset metadata so a customer can trace which integration produced a downstream result. Useful when handoff requires more than a diagram of the intended pipeline.
Field note: A standard and instrumentation ecosystem, not an automatic lineage scanner. Coverage depends on the integrations and facets you emit. -
lakeFS object-data rollbacksource-available
Branch and version object-store datasets without copying every object's bytes. Test a customer data transformation on a branch, then retain a commit that can be compared or reverted.
Field note: Works for data routed through lakeFS, not arbitrary OLTP writes. Retention and garbage collection determine which old objects remain recoverable. -
SQLMesh isolated change and backfill plans
Review SQL model changes as a plan against a named environment, then test affected models and backfill ranges in isolation. Useful when a warehouse change needs a controlled cutover.
Field note: Fits projects that use SQLMesh, not an add-on to any dbt project. Review engine-specific isolation and source snapshots before treating a preview as production evidence. -
SDV synthetic customer fixturessource-available
Generate tabular fixtures that reflect a customer's distributions and relationships without directly shipping the original rows. Useful when realistic test cases cannot use raw customer extracts.
Field note: Synthetic does not mean private. Review disclosure risk, business constraints, and the edition's supported models before sharing generated fixtures.
Expired credentials, partial responses, rate limits, and unavailable test windows belong in the acceptance suite.
-
Microcks multi-protocol contract mocks
Turn customer API artifacts into mocks and conformance checks across protocols including REST, SOAP, and asynchronous APIs. Useful before the real dependency or its sandbox is available.
Field note: Mock fidelity depends on the specification and examples. Confirm the business behavior against the real service during the customer test window. -
WireMock explicit failure scenarios
Program dependency responses and stateful sequences such as an expired token followed by successful reconnect. Useful when the acceptance test needs a precise failure, not a random outage.
Field note: Fixtures describe intended behavior. Track which pagination, rate-limit, and error cases were confirmed with the real API. -
Hoverfly capture then simulate
Capture HTTP interactions through a proxy, export the simulation, and replay it without repeatedly calling a customer dependency. Stateful simulations preserve meaningful request sequences.
Field note: Capture needs the client traffic to pass through the proxy. Scrub credentials and customer payloads before sharing simulations. -
Keploy application-call replay
Record application traffic and dependency interactions to replay integration tests with captured mocks. Useful when a working customer path exists but its regression fixtures do not.
Field note: Check the capture mode, runtime support, and ability to instrument the caller. Recorded traffic can contain sensitive customer data. -
Toxiproxy network fault injection
Inject latency, disconnects, and bandwidth constraints into TCP dependencies. Prove that a connector resumes or fails predictably when the network becomes unreliable.
Field note: Tests transport failure, not application-level permission errors or malformed JSON. Add those cases separately. -
Schemathesis schema-driven edge cases
Generate property-based API tests from OpenAPI or GraphQL schemas. Find invalid inputs and response-contract failures beyond the successful requests used in the demo.
Field note: Use an authorized test environment. Add explicit assertions for business rules and writes that the schema cannot describe. -
oasdiff breaking API changes
Compare OpenAPI revisions and report breaking changes before a customer integration is updated. Useful when an upstream team changes fields, required parameters, or response contracts.
Field note: Compares declared specifications. It cannot detect an undocumented change in the live API. -
mitmproxy inspect actual traffic
Intercept and modify authorized HTTP traffic to diagnose undocumented headers, redirects, and payload behavior. Useful when the observed customer integration contradicts its documentation.
Field note: Requires permission and a compatible trust setup. TLS pinning and locked customer clients can prevent interception. -
Pact consumer-driven contracts
Capture the interactions a customer's consuming application relies on, then verify them against the provider. Useful when different customers depend on different parts of the same API.
Field note: Choose the language binding for the application. Contract verification is not provider functional testing or an end-to-end business acceptance test. -
Mountebank raw-protocol simulations
Stub text or binary TCP exchanges when a customer vendor interface is not HTTP. SMTP capture can verify that the integration sent the expected message.
Field note: TCP fixtures need explicit message framing; SMTP capture does not stub responses. The repository documents a maintainer transition, so review required protocol support.
Keep the source evidence and test the failure cases before treating extracted content as a business record.
-
Docling layout-aware parsing
Convert customer documents into structured content with reading order and table information. Useful when plain-text extraction loses the columns or layout that carry the meaning.
Field note: Inspect representative files against the originals. OCR and layout models add runtime dependencies; successful conversion does not guarantee accurate fields. -
DocETL document-to-record pipelines
Define LLM-based map, reduce, resolve, and related operations over document collections. Useful for extracting domain-specific records when a format parser cannot express the task.
Field note: Different job from PDF parsing. Retain labeled examples and inspect the pipeline; generated validators and model outputs still need review. -
Presidio PII detection and masking
Detect and anonymize sensitive text before approved logs, fixtures, or model calls leave the customer environment. Add domain-specific recognizers for identifiers the defaults miss.
Field note: Detection has false positives and false negatives. Test customer-specific formats; a clean scan is not proof that a payload contains no sensitive data. -
pdfplumber geometry-based PDF tables
Inspect characters, ruling lines, and cell boundaries with visual debugging. Useful for a customer's repeatable invoice or report layout where explicit geometry is preferable to a layout model.
Field note: Works best on machine-generated PDFs. It does not add OCR to scans; retain failing examples when tuning table rules. -
docTR controllable local OCR
Choose text-detection and recognition models for customer images and scans. Useful when OCR must stay inside a restricted environment and the parsing pipeline needs model-level control.
Field note: The current library uses PyTorch. Stage weights and dependencies locally, and test scan quality; OCR output is not a validated business record. -
PaddleOCR multilingual document recognition
Run OCR and structured document parsing over customer scans with language- and layout-oriented models. Useful for mixed-language forms, tables, and formulas that need local processing.
Field note: Choose and test the specific model, language coverage, and hardware requirement. Lightweight OCR and vision-language variants have different operating costs. -
OCRmyPDF searchable scanned PDFs
Add an OCR text layer to scanned PDFs so existing page-based records can be searched and copied. Useful when the deliverable must remain a PDF rather than become extracted JSON.
Field note: Stage OCR language data for isolated installs. Signed PDFs are refused by default; modifying an approved copy can invalidate its signature. This is not PDF malware sanitization.
Persist the business process and design for the crash after a downstream write but before its acknowledgement.
-
Temporal long-running business workflows
Persist workflows through crashes and long waits for customer approval. Signals, activities, and execution history help coordinate business steps across unreliable systems.
Field note: Fits a customer who can operate or approve the workflow service. Retried external activities still require idempotency or reconciliation. -
Restate keyed durable handlerssource-available
Journal durable handlers and keep state around customer-specific entities. Useful for serializing work on the same record and resuming multi-step integration calls after failure.
Field note: Has its own runtime. An external side effect can repeat before its result is journaled; invocation deduplication alone does not prevent that. -
DBOS Postgres-backed workflows
Add durable workflows and queues through a library that checkpoints into PostgreSQL. Useful when the customer already has Postgres and an extra orchestration service is difficult to operate.
Field note: Completed steps are checkpointed; unfinished external calls may repeat. Exactly-once database transactions do not imply exactly-once API writes. -
Windmill operator approval pages
Suspend a script flow and expose approval or cancellation through generated resume URLs and a run interface. Useful when the customer's operator needs an explicit human gate around an integration step.
Field note: Treat resume URLs as bearer credentials and enforce the required approver checks. A visible run history is not a tamper-evident approval ledger. -
Hatchet per-customer execution budgets
Use customer-keyed concurrency, fair scheduling, and rate limits to keep one backfill from monopolizing workers or exhausting an upstream API budget.
Field note: Adds an orchestration service. Configure keys, queue policy, and failure handling; cancellation cannot undo a remote write already committed. -
LangGraph persisted agent interruption
Checkpoint an agent workflow and interrupt before a tool action for customer review. Resume the saved thread with an explicit decision rather than asking the model to remember an approval.
Field note: Production needs a persistent checkpointer. Resume restarts the node, so earlier side effects must be idempotent. Supply authorization and an approval record separately.
Test direct tool behavior, then the model's choices, then the resulting business state.
-
MCP Inspector direct protocol testing
Exercise a customer MCP server's tools, resources, and prompts without an agent deciding what to call. Separate transport, authentication, and argument failures from model behavior.
Field note: Protocol success is not business authorization. Test denied access and actual side effects, including tools advertised as read-only. -
Snyk agent-scan agent toolchain audit
Scan MCP servers, agent configurations, and skills for excessive permissions, tool poisoning, and injection exposure before they touch customer credentials. Useful when a community MCP server for the customer's system needs a security answer before connection.
Field note: Covers known patterns, not the server's runtime behavior. Treat a clean scan as one input to the customer's security review, and re-scan on version changes. -
promptfoo customer-case regressions
Version customer examples and compare prompts or models with assertions and adversarial cases. Useful for turning acceptance examples into a rerunnable regression suite.
Field note: Keep a held-out case set. A text or model-graded answer can pass while the downstream record is wrong; check resulting state separately. -
Inspect AI task-level evaluation
Define tasks, solvers, and scorers with execution logs for multi-step agent behavior. Useful when acceptance depends on completing a customer workflow rather than producing a plausible answer.
Field note: Supply the real task fixtures and business-state scorer. An evaluation framework cannot decide what the customer's successful outcome means. -
Langfuse production-case traces
Capture model, retrieval, and tool-call traces, attach scores, and turn reviewed customer cases into datasets. Useful for connecting acceptance regressions to failures observed in the delivered application.
Field note: Control which customer payloads enter telemetry. Trace visibility and replay do not prove the final business state; deployment features vary by edition. -
Phoenix retrieval and trace experimentssource-available
Inspect OpenTelemetry-based AI traces and compare retrieval or response experiments over versioned examples. Useful for distinguishing an extraction or retrieval failure from a prompt change.
Field note: Configure approved telemetry storage and model endpoints. Evaluator scores need calibration; the source-available project is separate from the managed Arize platform. -
DeepEval Python task-regression checks
Add LLM and agent metrics to Python test suites alongside customer fixtures. Useful when domain regression checks should run with the implementation's existing test code.
Field note: Many metrics use model judges and incur calls. Keep deterministic business-state assertions and validate graders against examples reviewed by the customer. -
garak model vulnerability probes
Run selected probe and detector families against an approved model endpoint for injection, leakage, and other failure modes. Useful before exposing a customer-facing model workflow.
Field note: Scope the scan and call budget to a test environment. Detected failures need investigation; a clean scan is not a safety or authorization proof. -
PyRIT multi-turn adversarial testing
Orchestrate multi-turn red-team strategies against an authorized AI target. Useful for testing whether an agent can be steered away from the customer's approved behavior over several exchanges.
Field note: Use approved test data and targets, with request budgets and review. Attack outcomes do not replace tool-level access checks or a business acceptance suite.
| Read | Why it belongs |
|---|---|
| Customer discovery, embedded engineering, and deciding what field work should become a product capability. | |
| An operational model that connects data, decisions, actions, and security. Foundry-specific context, not a required platform. | |
| Tasks, trials, graders, and checking what happened after an agent acted. | |
| FDE learning, practice write-ups, and case material. A sibling list with a broader mandate than this catalog. | |
| Principles for production-grade agent software; the discipline behind several entries here. | |
| Practical patterns and code for taking agents from demo to production. |
Use the delivery recipes to plan an acceptance test, and the engagement brief to agree on the customer outcome, constraints, owner, and product follow-up.
Research proceeds in loops: discover candidates with Grok, compare neighboring tools, challenge claims, check primary sources, and cut entries that do not add enough. The research record includes decisions and corrections. Listed capabilities come from documentation, not a claim that every tool has been deployed in a customer environment.
Most entries are open source. source-available marks restricted projects; verify the deployment edition you need. Icons identify the maintaining GitHub account or the project's own mark. Website links use favicons.
Contributions should name the customer failure, show the documented mechanism that addresses it, and explain why the nearest listed alternative is insufficient. Read CONTRIBUTING.md.