Deploy Nexus BYOC
Install and operate Pinecone Nexus in your own cloud account.
For architecture, see the Nexus BYOC overview. For data residency and limits, see Data residency and limits.
Prerequisites
Section titled “Prerequisites”Before deploying Nexus BYOC, ensure you have the following tools installed on the machine that runs the install:
| Tool | Purpose | Install |
|---|---|---|
| Git | Clone the deployment repository | git-scm.com |
| Python 3.12+ | Runtime | python.org |
| uv | Package manager | docs.astral.sh/uv |
| Pulumi | Infrastructure-as-code | pulumi.com/docs/install |
| kubectl | Cluster access | kubernetes.io |
You also need:
- The CLI for your cloud provider:
- AWS: AWS CLI v2
- GCP: gcloud CLI, plus the
gke-gcloud-auth-plugincomponent (gcloud components install gke-gcloud-auth-plugin) - Azure: Azure CLI
- A dedicated cloud account with admin-level permissions, used only for this deployment:
- AWS: a dedicated account with
AdministratorAccess. The installer creates IAM roles and policies, soPowerUserAccessisn't sufficient. - GCP: a dedicated project with
roles/ownerand billing enabled. The installer creates IAM service accounts and bindings, soroles/editorisn't sufficient. - Azure: a dedicated subscription with
Owner. The installer creates managed identities and role assignments, soContributorisn't sufficient.
- AWS: a dedicated account with
- A Pinecone API key from the Pinecone console.
- A Pinecone Enterprise plan (required for BYOC access).
- A generation-LLM key (BYOM). The shipped default is Google Gemini, so get a key from Google AI Studio. Embedding and rerank run on Pinecone-hosted models by default (no extra key). See Model guidance for the model catalog and how to repoint generation, embedding, or rerank, and Data flows and residency for where each call's content goes.
- A Pulumi state backend (either Pulumi Cloud or local state via
pulumi login --local). - Sufficient cloud quota for the resources (the setup wizard validates this).
Confirm these environment inputs before you start:
- Region and three availability zones (AZs). Deploy across three AZs, the supported high-availability shape. The generation-model provider you choose must be available in the region.
- A spare private IP range (RFC 1918,
/16to/20) that doesn't overlap your existing networks. This becomes the Nexus virtual network. - Where your source data lives. Know where your corpus resides so you can stage it into a context after install.
- Egress paths. Nexus makes outbound connections for the control-plane callback, metrics and traces, container image pulls, and calls to the inference models you configure. What content those model calls carry depends on how you configure your models. See Data flows and residency.
1. Deploy
Section titled “1. Deploy”To deploy Nexus BYOC, follow these steps.
Authenticate
The setup script checks your credentials but doesn't log you in, so authenticate to your cloud and to Pulumi first.
Bash aws configure # or: aws sso login / exported AWS_* env vars aws sts get-caller-identity # verify pulumi login # or: pulumi login --localBash gcloud auth login gcloud auth application-default login pulumi login # or: pulumi login --localBash az login az account show # verify pulumi login # or: pulumi login --localIf you use the local Pulumi backend, choose a passphrase for encrypting stack secrets and export it as
PULUMI_CONFIG_PASSPHRASE. Everypulumicommand needs it.Run the setup wizard
Clone the deployment repository and run the bootstrap script from the clone. The generated project is created next to the clone and depends on it.
Bash git clone https://github.com/pinecone-io/pulumi-pinecone-nexus-byoc.git bash pulumi-pinecone-nexus-byoc/bootstrap.sh --cloud gcpUse
--cloud awsor--cloud azurefor the other clouds, and--stack-name <name>to name the Pulumi stack (default:prod).The script selects your cloud provider, checks that required tools are installed, verifies your cloud credentials, prompts for the project directory and name, then launches an interactive wizard that collects your configuration, validates your quotas, and generates a Pulumi project in an adjacent directory (default:
pinecone-nexus-byoc). No cloud resources are created during this step.Setup wizard prompts
The wizard prompts you for the following:
Prompt Description Cloud provider AWS, GCP, or Azure (skipped if pre-selected via --cloud).Project directory and name Where to generate the Pulumi project. Defaults to pinecone-nexus-byoc.Pinecone API key Your API key from the Pinecone console (or uses PINECONE_API_KEY).Cloud credentials Validates credentials and displays your account/project/subscription ID. Region Region for deployment. Availability zones Three zones for high availability. VPC/VNet CIDR block Private IP range for the deployment. Choose a /16to/20that doesn't overlap existing networks.Generation-LLM key The default catalog's Gemini API key (BYOM). Network access Public access. Private-only access is not supported for Nexus. See Limitations. Deletion protection Whether to protect storage and database resources from accidental deletion. Enable to guard against teardown mistakes. Disable before pulumi destroy.Preflight checks Validates cloud quotas. If checks fail, request quota increases before proceeding. Pulumi backend Local ( ~/.pulumiwith passphrase) or Pulumi Cloud.After completing the wizard, a Pulumi project is generated in your project directory. To change configuration later, edit
Pulumi.<stack>.yamland runpulumi up.Deploy the infrastructure
Deploy the generated Pulumi project to create your cloud resources:
Bash cd pinecone-nexus-byoc pulumi upPulumi shows a preview of all resources to be created. Confirm to proceed. Provisioning time depends on the cloud:
Cloud Typical provisioning time GCP 25 to 30 minutes AWS 25 to 40 minutes Azure about 35 minutes When complete, the output displays:
- The
update_kubeconfig_commandfor configuring cluster access. - Your BYOC environment name.
- Two workspace console links, printed once the first-run
defaultworkspace reachesReady:nexus_default_workspace_data_console_url: the workspace console for thedefaultworkspace, served from your deployment (where you work with contexts and run queries).nexus_default_workspace_control_console_url: the Pinecone Console page for that workspace.
Infrastructure provisioned
The deployment creates the following in your cloud account:
Component AWS GCP Azure VPC / Networking VPC, public and private subnets, NAT gateways, internet gateway VPC network, subnets, Cloud NAT, Cloud Router VNet, subnets, NAT gateway Kubernetes EKS cluster with managed node groups GKE cluster with node pools AKS cluster with agent pools Object storage S3 buckets (corpus, knowledge artifacts, data, WAL, backups) GCS buckets Blob Storage containers Metadata store FoundationDB (Nexus metadata) FoundationDB FoundationDB Load balancing Network Load Balancer Internal load balancer Internal load balancer DNS Route 53 hosted zone Cloud DNS managed zone Azure DNS zone TLS certificates AWS Certificate Manager cert-manager cert-manager IAM IAM roles and policies Service accounts and Workload Identity Managed identities and Workload Identity The cluster comes up small and then autoscales across several node pools spread over the three AZs. See Cluster footprint for the node pools.
- The
Verify the deployment
Configure
kubectlusing theupdate_kubeconfig_commandfrom the deployment output:Bash aws eks update-kubeconfig --region <region> --name <cluster-name>Bash gcloud container clusters get-credentials <cluster-name> --region <region> --project <project-id>Bash az aks get-credentials --resource-group <resource-group> --name <cluster-name>Cluster access is for administrative tasks like viewing operations and troubleshooting. Everyday work (creating contexts, curating sources, and running queries) uses the Nexus console, CLI, or API.
Verify all components are running:
Bash kubectl get pods -A | grep -E "(pinecone|pc-|nexus)"All pods should show
Runningstatus. If any are inPendingorCrashLoopBackOff, see Troubleshooting.
2. Use
Section titled “2. Use”Once your deployment is up and the default workspace is ready, you work with Nexus the same way you would in the managed service. For the end-to-end lifecycle (create a context, stage sources, curate, and query with KnowQL), see the Nexus quickstart.
Two things are specific to BYOC:
-
Point your client at your own workspace host. Instead of the Pinecone-hosted endpoint, use your deployment's workspace host, the base of
nexus_default_workspace_data_console_urlfrom the deployment output. For the CLI, pass it as--api-url:Bash nexus login --api-url https://default-<vault>.wksp.<environment>.pinecone.io --api-key "$PINECONE_API_KEY"Authentication uses your Pinecone API key, and the tenancy boundary is the workspace's Pinecone project.
-
A
defaultworkspace is created on the firstpulumi up. The install creates it automatically and prints its data console URL (served from your deployment) and its Pinecone Console URL. Laterpulumi upruns never recreate or modify it.
3. Manage
Section titled “3. Manage”Operations and upgrades
Section titled “Operations and upgrades”Pinecone uses a pull-based model for cluster operations:
- When upgrades, scaling, or maintenance are needed, Pinecone queues operations in the control plane.
- An agent running in your cluster (deployed automatically during setup) continuously pulls pending operations.
- Operations execute locally within your cluster.
- Status is reported back to Pinecone for monitoring.
This model ensures Pinecone never needs direct access to your infrastructure. All communication is outbound from your cluster.
A deployment is pinned to two independent image tags that roll separately: pinecone-version (the Pinecone Database images) and nexus-version (the Nexus images). The two pins are unrelated. Bumping one doesn't touch the other. Pinecone manages upgrades in the background. To trigger one manually, set either pin (or both) to your target version (for example, main-abc1234) and re-run pulumi up:
# Bump the Pinecone Database version
pulumi config set pinecone-version <new-db-tag>
pulumi up
# Bump the Nexus version
pulumi config set nexus-version <new-nexus-tag>
pulumi upMonitoring
Section titled “Monitoring”You can monitor your deployment through multiple channels:
Pinecone console
View workspace and index metrics in the Pinecone console. Control plane operations and metrics work regardless of your network access mode.
Prometheus
To use Prometheus, configure your monitoring tool within your VPC to scrape metrics from the cluster. Your Prometheus instance must have network access to the BYOC VPC. The deployment output includes the metrics endpoint URL and port.
Audit logs
Cluster operations are persisted as Kubernetes CRDs for compliance and auditing:
kubectl get cluster-operationsCleanup
Section titled “Cleanup”To destroy your deployment:
# 1. Delete all workspaces via the Pinecone console
# 2. Then destroy the infrastructure
pulumi destroyIf deletion-protection is enabled (the default), you must either disable it in Pulumi.<stack>.yaml and run pulumi up, or manually delete the protected storage and database resources via the cloud console before running pulumi destroy.
Troubleshooting
Section titled “Troubleshooting”Preflight check failures
The setup wizard validates cloud quotas before deployment. If checks fail:
| Check | Resolution |
|---|---|
| VPC / network quota | Request a limit increase via your cloud provider's quota console |
| Kubernetes cluster quota | Request an EKS, GKE, or AKS cluster limit increase |
| IP address quota | Release unused IPs or request a limit increase |
| Instance / machine type availability | Verify the required type is available in your region |
| vCPU quota | Request a regional vCPU increase |
| Required APIs / providers | Enable the cloud APIs (GCP) or register the resource providers (Azure) the installer requires |
Deployment failures
If pulumi up fails partway through:
pulumi refresh # Sync state with actual resources
pulumi up # Retry deploymentOn AWS, the slowest single step is VPC endpoint service private DNS verification: roughly 15 minutes of Waiting for domain verification (pendingVerification) polling is normal, not a hang. Let it finish.
Cluster access issues
Ensure your cloud credentials match the account where the cluster is deployed:
aws sts get-caller-identitygcloud auth list
gcloud config get-value projectaz account showWorkspace stuck initializing
The first-run default workspace is created asynchronously and becomes ready only once the Nexus services in your deployment are up. If pulumi up times out waiting for it, re-run pulumi up once the cluster pods are Running. If it remains stuck, contact Pinecone support.
For additional help, see the GitHub Issues for the deployment repository.