Run PiyAPI your way
As a library
Install the lightweight SDK directly into your TypeScript or Python application for immediate, embedded bitemporal memory.
As a self-hosted server
Run the full PiyAPI Gateway and Memory Engine as an independent Docker service with complete control over Redis caching,
pgvector indexing, and LLM routing.Get started
Python Quickstart
Install the
piyapi-memory package and verify that hybrid memory storage and retrieval work in a few lines of code.Node.js Quickstart
Install the
@piyapi/sdk package and connect your first stateful agent workflow.Self-hosted server
Deploy the PiyAPI Docker container locally or on your Kubernetes cluster and configure your bitemporal infrastructure.
Go further
Configure components
Configure BYOK routing, embedding providers,
pgvector storage, and tune the UnifiedScorer for your specific deployment.Self-hosting features
Explore enterprise namespace isolation, semantic cache purging, multimodal CDC connectors, and compliance-grade redaction tools.
Build with cookbooks
Use practical examples for LangChain agents, CrewAI, MCP server integrations, and real-world compliance workflows.
Need a managed alternative? Compare the self-hosted open-source deployment with the fully managed PiyAPI Cloud Platform, or switch to the Platform documentation.
Default components
PiyAPI ships with sane, production-ready defaults for each part of the pipeline. Every component can be replaced or explicitly configured to match your existing VPC infrastructure and provider preferences.Library defaults
Self-hosted defaults
Configure your stack
Self-hosted architecture
Configuration
LLMs
LLMs
Route generation through any provider using our unified BYOK manager. Support for OpenAI, Anthropic, Gemini, DeepSeek, Mistral, Groq, Cohere, and Perplexity is handled natively.
Embeddings
Embeddings
Choose the embedding provider and model used for vector generation. Easily swap out defaults for specialized models depending on your latency and accuracy requirements.
Vector databases
Vector databases
Connect the storage backend used for dense HNSW retrieval. Defaults to PostgreSQL with
pgvector for robust, transaction-safe vector operations.Storage
Storage
Configure where persistent bitemporal graph data (PiyGraph) and conversational history are stored.
Metadata filtering
Metadata filtering
Control how stored context is filtered and retrieved using strict namespace scoping (
X-Namespace-Prefix) and arbitrary tagging rules.Reranking
Reranking
Fine-tune the UnifiedScorer reranking layer (combining BM25 trigram search and 11 distinct retrieval signals) when hyper-precise retrieval is required for your agents.
Build on the open-source stack
- Replace infrastructure components (for example, swap Redis for an alternative caching layer)
- Configure multi-provider routing and failover logic to prevent LLM downtime
- Integrate directly with your existing enterprise PostgreSQL databases
- Build custom Active Inference workflows and agent MCP tools
