OntoBricks Production Sizing Questionnaire

Use this questionnaire to collect the information needed to size a production OntoBricks deployment on Databricks. Complete all Required fields. Complete…

Use this questionnaire to collect the information needed to size a production OntoBricks deployment on Databricks. Complete all Required fields. Complete the Optional fields when the information is available; otherwise, state Unknown.

This questionnaire gathers workload characteristics. It does not constitute a price quote or capacity commitment. The resulting estimate should document all assumptions and include expected growth.

1. Customer and Databricks Environment

  • Required — Customer or project name: ______________________________
  • Required — Target go-live date: ______________________________
  • Required — Cloud provider: [ ] AWS [ ] Azure [ ] GCP
  • Required — Databricks region: ______________________________
  • Required — Databricks platform tier: ______________________________
  • Required — Unity Catalog enabled: [ ] Yes [ ] No
  • Required — Databricks Apps enabled: [ ] Yes [ ] No [ ] To confirm
  • Required — Lakebase Autoscaling available in the target region: [ ] Yes [ ] No [ ] To confirm
  • Optional — Compliance requirements: [ ] None [ ] HIPAA [ ] PCI [ ] Other: ______________________________
  • Optional — Enhanced Security and Compliance add-on required: [ ] Yes [ ] No [ ] To confirm

2. Users and Concurrency

  • Required — Total named OntoBricks users: ______________________________
  • Required — Expected concurrent users at peak: __________________________
  • Required — User distribution:
    • Viewers: ______________________________
    • Ontology editors: ______________________________
    • Graph builders or operators: ______________________________
    • Administrators: ______________________________
  • Required — Typical production usage window: ______ hours/day, ______ days/week
  • Optional — Expected concurrent API or MCP clients: _____________________
  • Optional — Expected annual user growth: ______________________________ %

3. Source Data

Repeat this section for each materially different source or workload.

Source or workload: ______________________________

  • Required — Source type: [ ] Unity Catalog table/view [ ] File [ ] External database [ ] SaaS application [ ] Other: ________________
  • Required — Current data volume: __________________ [ ] GB [ ] TB
  • Required — Number of tables or views: ______________________________
  • Required — Total row count: ______________________________
  • Required — Daily data change: __________________ [ ] GB [ ] TB [ ] rows
  • Required — Update pattern: [ ] Full refresh [ ] Incremental [ ] Change Data Capture [ ] Append-only
  • Required — Refresh frequency: ______________________________
  • Optional — Largest table: __________________ [ ] GB [ ] TB; __________________ rows
  • Optional — Average row width, if known: __________________ bytes
  • Optional — Data retention period: ______________________________
  • Optional — Existing ingestion service: ______________________________

Add another copy of this subsection for each additional source or workload.

4. Ontology and Mapping Complexity

  • Required — Number of business domains: ______________________________
  • Required — Number of production ontology versions retained per domain: ______________________________
  • Required — Estimated ontology size per domain:
    • Entity or class types: ______________________________
    • Relationship types: ______________________________
    • Attributes or properties: ______________________________
  • Required — Estimated number of source-to-ontology mappings: ______________________________
  • Required — Typical mapping complexity: [ ] Simple column mapping [ ] Joins across tables [ ] Complex SQL transformations [ ] Mixed
  • Optional — Maximum tables joined by one mapping: _______________________
  • Optional — Reasoning features used: [ ] None [ ] OWL 2 RL [ ] SWRL rules [ ] SHACL validation
  • Optional — Estimated number of SWRL rules: ____________________________
  • Optional — Estimated number of SHACL shapes: __________________________
  • Optional — Documents or binary artifacts stored per domain: __________________ [ ] GB [ ] TB

5. Knowledge Graph Builds

  • Required — Number of graph builds per day: ____________________________
  • Required — Build schedule: [ ] On demand [ ] Scheduled [ ] Both
  • Required — Maximum simultaneous builds: ______________________________
  • Required — Expected triples generated by a full build: __________________ [ ] million [ ] billion [ ] Unknown
  • Required — Maximum acceptable full-build duration: _____________________
  • Required — Build mode: [ ] Full rebuild [ ] Incremental [ ] Both [ ] To confirm
  • Optional — Expected triples changed per incremental build: __________________ [ ] million [ ] billion
  • Optional — Largest expected graph: __________________ [ ] million [ ] billion triples
  • Optional — Reasoning or validation executed during each build: [ ] None [ ] OWL 2 RL [ ] SWRL [ ] SHACL
  • Optional — Build-time availability requirement: [ ] Queries may pause [ ] Existing graph must remain queryable

If triple counts are unknown, provide representative source row counts and the expected number of entities, relationships, and attributes generated per row.

6. Graph Consumption

  • Required — Consumption methods: [ ] OntoBricks Graph Viewer [ ] GraphQL API [ ] MCP [ ] SQL [ ] Other: __________________________
  • Required — Peak interactive query rate: __________________ queries/second
  • Required — Typical query rate: __________________ queries/second
  • Required — Expected peak concurrent queries: __________________________
  • Required — Typical query complexity: [ ] Point lookup [ ] One-hop traversal [ ] Multi-hop traversal [ ] Aggregation [ ] Mixed
  • Required — Maximum acceptable interactive response time: ______________________________
  • Optional — Typical result size: __________________ rows or objects
  • Optional — Maximum traversal depth: __________________ hops
  • Optional — Bulk exports: ______ exports/day of approximately __________________ [ ] MB [ ] GB each
  • Optional — API traffic growth expected during the next 12 months: __________________ %

7. Lakebase Workload

OntoBricks uses Lakebase Autoscaling for its registry and graph storage.

  • Required — Estimated graph storage at go-live: __________________ [ ] GB [ ] TB [ ] Unknown
  • Required — Estimated graph storage after 12 months: __________________ [ ] GB [ ] TB [ ] Unknown
  • Required — Peak Lakebase query rate: __________________ queries/second [ ] Unknown
  • Required — Workload balance: ______ % reads / ______ % writes
  • Required — Required daily availability: __________________ hours/day
  • Optional — Required rollback or recovery retention: __________________ days
  • Optional — Largest expected write burst: __________________ rows/second
  • Optional — Known connection limit or connection-pooling constraints: ______________________________

8. SQL Warehouse Workload

OntoBricks uses a Databricks SQL Warehouse for Unity Catalog queries and Delta-backed operations.

  • Required — Existing SQL Warehouse available: [ ] Yes [ ] No [ ] To be provisioned
  • Required — Warehouse may be shared with other workloads: [ ] Yes [ ] No
  • Required — Expected OntoBricks warehouse activity: ______ hours/day, ______ days/week
  • Required — Maximum concurrent OntoBricks SQL statements: ______________________________
  • Required — Largest data volume scanned by one operation: __________________ [ ] GB [ ] TB [ ] Unknown
  • Optional — Preferred warehouse type: [ ] Serverless [ ] Pro [ ] Classic [ ] No preference
  • Optional — Existing warehouse size and scaling range: ______________________________
  • Optional — Auto-stop requirement: __________________ minutes

9. Databricks Apps and MCP

A standard deployment contains the OntoBricks application and an optional MCP companion application.

  • Required — MCP server required: [ ] Yes [ ] No [ ] To confirm
  • Required — Expected app usage: ______ hours/day, ______ days/week
  • Required — Peak simultaneous browser sessions: ________________________
  • Optional — Peak simultaneous MCP sessions: ____________________________
  • Optional — App compute preference: [ ] Medium [ ] Large [ ] Let the sizing team recommend
  • Optional — High availability or minimum replica requirement: ______________________________
  • Optional — External systems calling OntoBricks APIs: ______________________________

10. AI-Assisted Features

Complete this section only if OntoBricks AI-assisted design or automation features will be enabled. Foundation Model API consumption must be estimated separately from the core Databricks compute estimate.

  • Required — AI-assisted features enabled: [ ] Yes [ ] No
  • If yes — Features used: [ ] Ontology generation [ ] Mapping generation [ ] Automated assignment [ ] Other: ______________________________
  • If yes — Expected requests per day: ______________________________
  • If yes — Peak simultaneous requests: ______________________________
  • Optional — Typical input size: __________________ tokens or __________________ characters
  • Optional — Typical output size: __________________ tokens
  • Optional — Required model or serving endpoint: ________________________
  • Optional — Data residency or model governance constraints: ______________________________

11. Service Levels, Security, and Growth

  • Required — Production availability target: ____________________________
  • Required — Recovery time objective (RTO): _____________________________
  • Required — Recovery point objective (RPO): ____________________________
  • Required — Expected data growth over 12 months: _____________________ %
  • Required — Expected query growth over 12 months: ____________________ %
  • Required — Expected graph-build growth over 12 months: ______________ %
  • Optional — Network requirements: [ ] Public workspace connectivity [ ] Private Link [ ] IP access lists [ ] Other: ______________________
  • Optional — Customer-managed keys required: [ ] Yes [ ] No
  • Optional — Audit-log retention requirement: ___________________________
  • Optional — Planned traffic peaks or seasonal events: ______________________________

12. Representative Performance Sample

When possible, provide a representative production sample or benchmark:

  • Source data volume: ______________________________
  • Source row count: ______________________________
  • Ontology and mapping used: ______________________________
  • Generated triple count: ______________________________
  • Full-build duration: ______________________________
  • Incremental-build duration: ______________________________
  • Peak memory or compute observed: ______________________________
  • Typical query and measured latency: ______________________________
  • Environment where the measurement was taken: ____________________________

13. Additional Context

  • Current solution being replaced or complemented: ______________________________
  • Known bottlenecks or performance concerns: ______________________________
  • Constraints not covered above: ______________________________
  • Additional comments: ______________________________

14. Sizing Summary — For the Sizing Team

The customer may leave this section blank.

  • Recommended Databricks App compute: ______________________________
  • Recommended MCP App compute: ______________________________
  • Recommended Lakebase capacity and storage: _____________________________
  • Recommended SQL Warehouse type, size, and scaling range: ______________________________
  • Estimated monthly workload by component: ______________________________
  • Foundation Model API estimate, if applicable: __________________________
  • Growth headroom applied: ______________________________ %
  • Key assumptions: ______________________________
  • Exclusions: ______________________________
  • Confidence level: [ ] Low [ ] Moderate [ ] High
  • Reassessment trigger: ______________________________