OntoBricks Production Sizing Questionnaire
Use this questionnaire to collect the information needed to size a production OntoBricks deployment on Databricks. Complete all Required fields. Complete…
Use this questionnaire to collect the information needed to size a production OntoBricks deployment on Databricks. Complete all Required fields. Complete the Optional fields when the information is available; otherwise, state Unknown.
This questionnaire gathers workload characteristics. It does not constitute a price quote or capacity commitment. The resulting estimate should document all assumptions and include expected growth.
1. Customer and Databricks Environment
- Required — Customer or project name: ______________________________
- Required — Target go-live date: ______________________________
- Required — Cloud provider: [ ] AWS [ ] Azure [ ] GCP
- Required — Databricks region: ______________________________
- Required — Databricks platform tier: ______________________________
- Required — Unity Catalog enabled: [ ] Yes [ ] No
- Required — Databricks Apps enabled: [ ] Yes [ ] No [ ] To confirm
- Required — Lakebase Autoscaling available in the target region: [ ] Yes [ ] No [ ] To confirm
- Optional — Compliance requirements: [ ] None [ ] HIPAA [ ] PCI [ ] Other: ______________________________
- Optional — Enhanced Security and Compliance add-on required: [ ] Yes [ ] No [ ] To confirm
2. Users and Concurrency
- Required — Total named OntoBricks users: ______________________________
- Required — Expected concurrent users at peak: __________________________
- Required — User distribution:
- Viewers: ______________________________
- Ontology editors: ______________________________
- Graph builders or operators: ______________________________
- Administrators: ______________________________
- Required — Typical production usage window: ______ hours/day, ______ days/week
- Optional — Expected concurrent API or MCP clients: _____________________
- Optional — Expected annual user growth: ______________________________ %
3. Source Data
Repeat this section for each materially different source or workload.
Source or workload: ______________________________
- Required — Source type: [ ] Unity Catalog table/view [ ] File [ ] External database [ ] SaaS application [ ] Other: ________________
- Required — Current data volume: __________________ [ ] GB [ ] TB
- Required — Number of tables or views: ______________________________
- Required — Total row count: ______________________________
- Required — Daily data change: __________________ [ ] GB [ ] TB [ ] rows
- Required — Update pattern: [ ] Full refresh [ ] Incremental [ ] Change Data Capture [ ] Append-only
- Required — Refresh frequency: ______________________________
- Optional — Largest table: __________________ [ ] GB [ ] TB; __________________ rows
- Optional — Average row width, if known: __________________ bytes
- Optional — Data retention period: ______________________________
- Optional — Existing ingestion service: ______________________________
Add another copy of this subsection for each additional source or workload.
4. Ontology and Mapping Complexity
- Required — Number of business domains: ______________________________
- Required — Number of production ontology versions retained per domain:
______________________________ - Required — Estimated ontology size per domain:
- Entity or class types: ______________________________
- Relationship types: ______________________________
- Attributes or properties: ______________________________
- Required — Estimated number of source-to-ontology mappings:
______________________________ - Required — Typical mapping complexity: [ ] Simple column mapping [ ] Joins across tables [ ] Complex SQL transformations [ ] Mixed
- Optional — Maximum tables joined by one mapping: _______________________
- Optional — Reasoning features used: [ ] None [ ] OWL 2 RL [ ] SWRL rules [ ] SHACL validation
- Optional — Estimated number of SWRL rules: ____________________________
- Optional — Estimated number of SHACL shapes: __________________________
- Optional — Documents or binary artifacts stored per domain: __________________ [ ] GB [ ] TB
5. Knowledge Graph Builds
- Required — Number of graph builds per day: ____________________________
- Required — Build schedule: [ ] On demand [ ] Scheduled [ ] Both
- Required — Maximum simultaneous builds: ______________________________
- Required — Expected triples generated by a full build: __________________ [ ] million [ ] billion [ ] Unknown
- Required — Maximum acceptable full-build duration: _____________________
- Required — Build mode: [ ] Full rebuild [ ] Incremental [ ] Both [ ] To confirm
- Optional — Expected triples changed per incremental build: __________________ [ ] million [ ] billion
- Optional — Largest expected graph: __________________ [ ] million [ ] billion triples
- Optional — Reasoning or validation executed during each build: [ ] None [ ] OWL 2 RL [ ] SWRL [ ] SHACL
- Optional — Build-time availability requirement: [ ] Queries may pause [ ] Existing graph must remain queryable
If triple counts are unknown, provide representative source row counts and the expected number of entities, relationships, and attributes generated per row.
6. Graph Consumption
- Required — Consumption methods: [ ] OntoBricks Graph Viewer [ ] GraphQL API [ ] MCP [ ] SQL [ ] Other: __________________________
- Required — Peak interactive query rate: __________________ queries/second
- Required — Typical query rate: __________________ queries/second
- Required — Expected peak concurrent queries: __________________________
- Required — Typical query complexity: [ ] Point lookup [ ] One-hop traversal [ ] Multi-hop traversal [ ] Aggregation [ ] Mixed
- Required — Maximum acceptable interactive response time:
______________________________ - Optional — Typical result size: __________________ rows or objects
- Optional — Maximum traversal depth: __________________ hops
- Optional — Bulk exports: ______ exports/day of approximately __________________ [ ] MB [ ] GB each
- Optional — API traffic growth expected during the next 12 months: __________________ %
7. Lakebase Workload
OntoBricks uses Lakebase Autoscaling for its registry and graph storage.
- Required — Estimated graph storage at go-live: __________________ [ ] GB [ ] TB [ ] Unknown
- Required — Estimated graph storage after 12 months: __________________ [ ] GB [ ] TB [ ] Unknown
- Required — Peak Lakebase query rate: __________________ queries/second [ ] Unknown
- Required — Workload balance: ______ % reads / ______ % writes
- Required — Required daily availability: __________________ hours/day
- Optional — Required rollback or recovery retention: __________________ days
- Optional — Largest expected write burst: __________________ rows/second
- Optional — Known connection limit or connection-pooling constraints:
______________________________
8. SQL Warehouse Workload
OntoBricks uses a Databricks SQL Warehouse for Unity Catalog queries and Delta-backed operations.
- Required — Existing SQL Warehouse available: [ ] Yes [ ] No [ ] To be provisioned
- Required — Warehouse may be shared with other workloads: [ ] Yes [ ] No
- Required — Expected OntoBricks warehouse activity: ______ hours/day, ______ days/week
- Required — Maximum concurrent OntoBricks SQL statements:
______________________________ - Required — Largest data volume scanned by one operation: __________________ [ ] GB [ ] TB [ ] Unknown
- Optional — Preferred warehouse type: [ ] Serverless [ ] Pro [ ] Classic [ ] No preference
- Optional — Existing warehouse size and scaling range:
______________________________ - Optional — Auto-stop requirement: __________________ minutes
9. Databricks Apps and MCP
A standard deployment contains the OntoBricks application and an optional MCP companion application.
- Required — MCP server required: [ ] Yes [ ] No [ ] To confirm
- Required — Expected app usage: ______ hours/day, ______ days/week
- Required — Peak simultaneous browser sessions: ________________________
- Optional — Peak simultaneous MCP sessions: ____________________________
- Optional — App compute preference: [ ] Medium [ ] Large [ ] Let the sizing team recommend
- Optional — High availability or minimum replica requirement:
______________________________ - Optional — External systems calling OntoBricks APIs:
______________________________
10. AI-Assisted Features
Complete this section only if OntoBricks AI-assisted design or automation features will be enabled. Foundation Model API consumption must be estimated separately from the core Databricks compute estimate.
- Required — AI-assisted features enabled: [ ] Yes [ ] No
- If yes — Features used: [ ] Ontology generation [ ] Mapping generation [ ] Automated assignment [ ] Other: ______________________________
- If yes — Expected requests per day: ______________________________
- If yes — Peak simultaneous requests: ______________________________
- Optional — Typical input size: __________________ tokens or __________________ characters
- Optional — Typical output size: __________________ tokens
- Optional — Required model or serving endpoint: ________________________
- Optional — Data residency or model governance constraints:
______________________________
11. Service Levels, Security, and Growth
- Required — Production availability target: ____________________________
- Required — Recovery time objective (RTO): _____________________________
- Required — Recovery point objective (RPO): ____________________________
- Required — Expected data growth over 12 months: _____________________ %
- Required — Expected query growth over 12 months: ____________________ %
- Required — Expected graph-build growth over 12 months: ______________ %
- Optional — Network requirements: [ ] Public workspace connectivity [ ] Private Link [ ] IP access lists [ ] Other: ______________________
- Optional — Customer-managed keys required: [ ] Yes [ ] No
- Optional — Audit-log retention requirement: ___________________________
- Optional — Planned traffic peaks or seasonal events:
______________________________
12. Representative Performance Sample
When possible, provide a representative production sample or benchmark:
- Source data volume: ______________________________
- Source row count: ______________________________
- Ontology and mapping used: ______________________________
- Generated triple count: ______________________________
- Full-build duration: ______________________________
- Incremental-build duration: ______________________________
- Peak memory or compute observed: ______________________________
- Typical query and measured latency: ______________________________
- Environment where the measurement was taken: ____________________________
13. Additional Context
- Current solution being replaced or complemented:
______________________________ - Known bottlenecks or performance concerns:
______________________________ - Constraints not covered above:
______________________________ - Additional comments:
______________________________
14. Sizing Summary — For the Sizing Team
The customer may leave this section blank.
- Recommended Databricks App compute: ______________________________
- Recommended MCP App compute: ______________________________
- Recommended Lakebase capacity and storage: _____________________________
- Recommended SQL Warehouse type, size, and scaling range:
______________________________ - Estimated monthly workload by component: ______________________________
- Foundation Model API estimate, if applicable: __________________________
- Growth headroom applied: ______________________________ %
- Key assumptions: ______________________________
- Exclusions: ______________________________
- Confidence level: [ ] Low [ ] Moderate [ ] High
- Reassessment trigger: ______________________________