DataLab Clearance
Guidance on output rules, available treatments, practical examples and how to submit a clearance request
Roles and responsibilities
Output clearance is a shared responsibility. Researchers prepare and submit outputs that meet the rules, and the DataLab Clearance team confirms that the outputs meet the rules and that any required treatments have been applied.
Researcher responsibilities
- Produce outputs, applying the DataLab output rules.
- Apply the treatments when an output does not meet a requirement.
- Clearly describe each output and provide supporting evidence for assessment.
DataLab Clearance team responsibilities
- Assess submitted outputs against the rules, based on the evidence provided.
- Release outputs once verified to meet all applicable rules.
- Provide guidance on the rules, treatments and related enquiries.
The clearance team aims to clear outputs within two weeks. However, this depends on the resources available and the completeness of the submission. Submissions that clearly show the rules have been applied, supported by evidence and clear descriptions, can be reviewed quickly.
Output rules
The most common types of analysis and their relevant output rules are listed below. Other output types will be assessed based on similar rules.
| Output type | Applicable rules |
|---|---|
| Frequency tables (counts, percentages) | Rule of 10 Group disclosure rule |
| Magnitude statistics (means, sums, ratios) | Rule of 10 Dominance rules |
| Quantiles (percentiles, medians) | Quantile rules |
| Minimums, maximums, ranges | Quantile rules |
| Models including regressions | Model rules |
| Charts, graphs, plots | Chart rules |
| Code | Code rules |
| Microdata | Not appropriate for output |
| Synthetic microdata | Not appropriate for output |
Rule of 10
The rule of 10 refers to the minimum number of contributors required for each cell or statistic. These may be individual people, households or organisations, depending on the output.
If an output includes counts of services, apply the rule of 10 to the number of people who received those services, not to the number of services.
When to apply the rule
This rule applies to most outputs, including but not limited to:
- counts or frequencies
- means and sums
- percentages, proportions and ratios of counts or sums
- charts, graphs and plots
- weighted counts (the rule of 10 applies to the underlying unweighted counts).
Key requirements
- Each statistic must be derived from at least 10 people (for person data), 10 households (for household data) or 10 organisations (for business data).
- Counts that could be deduced from totals or related values must also meet the rule of 10.
- The rule of 10 should be applied to the numerator and denominator separately.
What to do if the rule breaks
If an output does not meet the rule of 10, apply a treatment. The treatment must sufficiently modify the data so that no reported or deducible value is based on fewer than 10 contributors. Evidence of treatment must be provided.
Recommended treatments include:
Combining categories
Aggregate small categories or simplify cross-tabulations so that all statistics in the output satisfy the rule of 10.
Suppression
- Suppress cells that fail the rule of 10. (This is primary suppression.)
- Suppress additional cells so that cells suppressed by primary suppression cannot be calculated. (This is secondary suppression or consequential suppression.)
Suppression is generally easy to apply to small tables. However, it can be complex for larger tables or outputs where there are many relationships between statistics in an output.
Rounding
- Suppress counts that fail the rule of 10.
- Round all remaining counts to the nearest 10, or randomly to a multiple of 5. (Note that total counts can be rounded independently of their components.)
- Calculate any proportions, percentages, rates or means using the rounded counts.
Rounding is useful for large tables and outputs of counts. It removes the need to apply secondary suppression.
Differencing
Differencing refers to deriving values by subtracting related cells. Differencing can occur between tables, between subpopulations or when comparing a total cell to the sum of its components. The output rules apply to values that can be calculated by differencing, as well as to values explicitly included in outputs.
When to apply the rule
Differencing risks can occur between nested populations, overlapping subgroups, and categories that include missing values.
Key requirements
- Differences in counts, percentages, proportions, and contributors for sums must meet the rule of 10.
- Differences in sums must also meet the dominance rules.
What to do if the rule breaks
If differences fail the rule of 10, treat this by combining categories, suppressing values or rounding.
Dominance rules
The dominance rules specify the maximum amount a person or organisation can contribute to a cell or statistic.
When to apply the rule
Check dominance when reporting:
- totals or sums
- means or averages
- other magnitude statistics
- transactional values (for example, number of doctor visits).
The dominance rules do not apply to counts of people, households or organisations.
Key requirements
- (1,50) rule: the largest contributor must not contribute more than 50% of the total.
- (2,67) rule: the two largest contributors together must not contribute more than 67% of the total.
- If the data includes both positive and negative values, use absolute values to check dominance. Compare the largest values and totals using their absolute size.
- If the output is weighted, compare the largest unweighted values with the weighted totals to check dominance.
What to do if the rule breaks
If either the (1,50) or (2,67) rule fails, apply a treatment.
Recommended treatments include:
Combining categories
Combine categories until the output passes both dominance rules.
Suppression
- Suppress the affected values.
- If the suppressed statistic can be derived using totals or other values, apply secondary suppression.
Rounding
Round values such that the dominance rules pass. For example, round to the nearest 100 or 1,000.
Group disclosure rule
Group disclosure happens when almost all individuals in a group share a characteristic. This makes it possible to accurately predict that an individual has that characteristic, solely from knowing that they belong to that group. Group disclosure may pose a risk, even if individuals seem to be protected under other rules.
Whether the group disclosure rule applies depends on how sensitive the characteristic and group are.
When to apply the rule
Check group disclosure for tabular statistics, including but not limited to:
- counts
- percentages and proportions
- totals of small and unique groups.
Key requirements
- No cell should contain more than 90% of the total for its row or column.
What to do if the rule breaks
Recommended treatments include:
Combining categories
Combine categories so that no cell contains more than 90% of the row or column total or so that the population groups are not sensitive.
Suppression
- Suppress the affected cell and all related statistics.
- If the suppressed value can be derived using totals or other values, apply secondary suppression.
Rounding
Round counts, percentages or proportions so that no cell contains more than 90% of the total.
Quantile rules
The quantile rules specify the minimum number of contributors required in a sample. This minimum requirement is different for each type of quantile.
The quantile rules apply to all types of quantiles, including but not limited to:
- medians
- percentiles
- interquartile ranges
- quantiles displayed on a box and whisker plot
- quantiles of the distribution of residuals from a regression.
Key requirements
- Provide the underlying (unweighted) counts of people, households or organisations in the sample when reporting quantiles.
- Each quantile must be derived from a sample that meets the minimum number of contributors for that quantile.
The minimum number of contributors required for a quantile is given by multiplying the number of groups into which the quantile partitions the sample by 5. The table below provides the minimum numbers of contributors required for some common quantiles.
| Quantile | Minimum contributors |
|---|---|
| Medians ( 0.50 ) | 10 |
| Quartiles ( 0.25, 0.5, 0.75 ) | 20 |
| Quintiles ( 0.2, 0.4, 0.6, 0.8 ) | 25 |
| Deciles ( 0.1, 0.2, 0.3, …, 0.9 ) | 50 |
| Vigintiles ( 0.05, 0.1, 0.15, …, 0.95 ) | 100 |
| Percentiles ( 0.01, 0.02, …, 0.99 ) | 500 |
What to do if the rule breaks
Recommended treatments include:
Changing quantile types
Choose a different type of quantile if the sample has fewer contributors than the required minimum number.
Suppression
Suppress any quantile values that do not have the required minimum number of contributors.
Minimums and maximums
Minimum and maximum values must meet the rule of 10.
Key requirements
Each minimum or maximum value must be the value of at least 10 people (for person data), 10 households (for household data) or 10 organisations (for business data).
What to do if the rule breaks
If the minimum or maximum fails the rule of 10, report alternative quantiles that meet the quantile rules, such as:
- 1st and 99th percentiles
- 5th and 95th percentiles (vigintiles)
- 10th and 90th percentiles (deciles).
Model rules
Degrees of freedom
The degrees of freedom rule refers to the minimum residual degrees of freedom required to output results from a model. Results may include coefficients, predictions, performance metrics and diagnostics. The residual degrees of freedom are calculated by subtracting the number of independent variables or other parameters estimated in the model from the number of observations.
Key requirements
All models must have at least 10 degrees of freedom.
What to do if the rule breaks
Recommended treatments include:
- reducing the number of variables or parameters in the model
- increasing the sample size.
Ordinary least squares regression
These rules only apply to ordinary least squares (OLS) regression. They do not apply to other regressions such as fixed effects or two-stage regressions.
Key requirements
- The R-squared must be less than or equal to 0.9.
- There should be at least one continuous independent variable.
What to do if the rule breaks
If the R-squared is greater than 0.9, suppress the intercept.
If all independent variables are categorical, options include:
- suppressing the intercept
- adding at least one continuous independent variable
- checking that the tabular means of the dependent variable meet the rule of 10 and dominance rules.
Survival curves
Key requirements
Each step change in the curve must represent at least 10 people, households or organisations.
What to do if the rule breaks
Recommended treatments include:
- combining steps
- rounding counts at all steps, so that step changes cannot be accurately calculated.
Tree-based models
Key requirements
Each terminal leaf node must contain at least 10 unique contributors.
What to do if the rule breaks
Recommended treatments include:
- suppressing the results
- setting the minimum number of contributors in each terminal leaf node to at least 10 through hyperparameter tuning.
Correlation coefficients
Key requirements
Correlation coefficients and Gini coefficients must be based on at least 10 contributors.
What to do if the rule breaks
Suppress the results.
Other models
For other models not listed above, demonstrate:
- all results are based on at least 10 unique people or organisations
- the model outputs do not reveal information about individual people, households or organisations.
Chart rules
All charts, graphs, plots and other data visualisations must meet the rules that apply to their underlying data.
Key requirements
- Provide the data used to create the chart.
- Include evidence that the data meets the relevant output rules.
- Charts must not display individual person-level or business-level data.
- All plotted points or elements must represent or be derived from at least 10 contributors.
What to do if the rule breaks
Treat the underlying data using the treatments recommended for the relevant output rule.
Code rules
Key requirements
- Code must not contain counts or other data, including in comments.
- Code must not contain IDs, including hashed IDs.
If requesting code in a Jupyter Notebook (IPYNB), SAS Enterprise Guide Project (EGP) or similar file, save a version with all results, data and logs removed.
Alternatively, provide the code in plain text format, such as a Python script (PY) or SAS program (SAS).
Logs cannot be cleared and should not be requested.
Data-specific rules
Additional rules apply to some data. Below are the most common types. If unsure whether additional rules apply, submit the output and the DataLab Clearance team will advise.
Minimum geography for health data
Outputs using Medicare Benefits Schedule, Pharmaceutical Benefits Scheme or Australian Immunisation Register data have a minimum geography level of Statistical Area 3 or equivalent (for example, Local Government Areas).
Secondary contributor rule
For datasets containing both person-level and business-level data, outputs have minimum contributor requirements at both levels. For example, employee-level output must represent at least 10 people to meet the primary count requirement. It must also include employees from at least 5 organisations. These minimum contributor requirements may differ depending on the data.
Request DataLab Clearance
All DataLab inputs, outputs and transfers require ABS clearance. Do not transcribe any information into or out of the DataLab, including data, code or notes.
Submit all DataLab clearance requests through the myDATA portal. For instructions on how to submit, see Clearance requests in myDATA. Below is what can be requested.
Input clearance
Request the input of aggregate data, concordances, supporting material, or code that you have written to your DataLab project.
Requirements
- Code files must not contain IDs, counts or other data.
- For aggregate data, provide the applicable copyright, licence and conditions of use.
- Provide the source of the reference material (for example, data item lists or concordance files).
Items that will not be cleared
- Names of people or businesses in code or data files
- Addresses or specific location coordinates (longitude and latitude)
- Large free‑text fields where names or other identifiers cannot be reliably checked
- Code from external sources (for example GitHub), code packages, or any code that includes executables
Output clearance
Request clearance to release outputs from your DataLab project. Submit an output request for any data, results or code that you want to use outside the DataLab. The ABS must clear all outputs before they can be accessed outside the DataLab.
Requirements
To support an efficient clearance process, ensure you:
- apply all relevant output rules and provide evidence that the rules have been met
- apply appropriate treatments if a rule has failed
- create well-organised, clearly labelled files
- request clearance for only what you require.
Items that will not be cleared
- Outputs that do not meet the DataLab output rules
- Outputs at unit record level
- Synthetic data at unit record level
Transfer clearance
Request approval to transfer a small number of code files and non-data files between DataLab projects. Ensure the files contain no IDs, counts or other data, including in any logs or comments.
Requirements
- Remove all IDs, counts and other data from code files, including from comments.
- Provide a description and context for each file, including the names of any data packages used.
- Define the population scope and variables.
- If the projects have access to different data packages, apply output rules and provide evidence if required.
DataLab Clearance enquiries
To enquire about:
- the DataLab clearance process
- application of clearance rules
- clearance requirements or guidance
Use the 'Clearance requests' tile in the myDATA portal to contact the DataLab Clearance team. Do not include any uncleared data from the DataLab in your enquiry. If your enquiry relates to a clearance request that you have submitted, include the request ID.