Databricks
Databricks is available to SEAD users as a non-standard product
What is Databricks
Databricks is a cloud-based Big Data processing platform which provides users with an integrated environment to collaborate on projects and offers a range of tools for data exploration, visualisation and analysis. Within the Databricks environment, users can:
- Build pipelines for streaming data processing.
- Build and run machine learning tools.
- Create interactive dashboards.
- Take advantage of scalable distributed computing capability.
Users will also have access to the Databricks Academy training subscription (an online library of Databricks training guides), in addition to instruction materials on how to setup the Databricks workspace provided in the Shared Library (L:) drive as well as the example notebooks available through the Databricks UI.
How to allocate a Databricks workspace to a project
To allocate a Databricks workspace to your project, you will need to submit a request to your SEAD administrator. Once your project is allocated a Databricks workspace, it can be accessed from within your Virtual Machine (VM) using the installed Edge or Firefox browsers.
Cost consumption
Databricks costs can be complex. Understanding how costs are calculated is an important consideration for users and projects wanting access to Databricks. Information about indicative Databricks pricing can be found at Azure Databricks Pricing.
As Databricks uses separate compute power, projects requesting access to Databricks should consider if they need to maintain their existing VM sizes. The option of scaling down the size of existing VMs provides users the opportunity to save on project costs.
What are the cluster policy arrangements?
Users can select from the following pre-configured cluster policy options:
| Instance | Server Purpose | Max Autoscale workers | vCPU(s) | RAM/ | Databricks Units |
|---|---|---|---|---|---|
| DS3 v2 | General Purpose | 5 | 4 | 14GB | 0.75 |
| D13 v2 | Memory optimised | 4 | 8 | 56GB | 2 |
| F16s v2 | Compute optimised | 4 | 16 | 32GB | 3 |
The ‘SEAD standard cluster policy’ is also available. The SEAD standard policy allows for additional flexibility when choosing a suitable instance type for your workload. This policy sets a maximum Databricks Unit (DBU) consumption limit of 32 DBU per hour. Users are expected to consider their cluster configuration carefully to ensure it is proportionate to workload needs and budget limitations.
To ensure the security and integrity of SEAD, partners will not have administrative access to the Databricks workspace and some usage restrictions may apply. Administration will be exclusively managed by the ABS.
Note: The ABS provides information on appropriate Databricks cost management to end users within the shared library (L: Drive).