Combined Ratio Insurance
How the combined ratio measures underwriting performance, what its loss and expense components include, and why calendar and accident year figures differ.
What it takes to anonymize claims data, how that differs from pseudonymization and masking, and why removing names is only the start of the job.
Anonymizing insurance claims data means transforming information about claims so that the people it relates to are no longer identifiable in the relevant context. The aim is to preserve useful information about losses, costs, and claims handling while preventing the data from revealing individual claimants or other people involved.
Removing names is only part of the task. A claim can also identify someone through its location, event date, supporting documents, or an unusual combination of details.
Insurance claims contain information collected for a specific operational purpose: investigating an event and deciding how the policy responds. A claims file may contain contact details, photographs, repair estimates, medical information, payment records, and correspondence.
Not every later use needs that level of detail. A team studying average settlement time by claim type may need dates and durations, but it usually does not need a claimant’s name or bank account number.
Anonymization changes the dataset so that people cannot be identified from the resulting information using means reasonably likely to be available. The assessment depends on context, including who receives the information and what other information they could use. The ICO explains that removing obvious identifiers alone does not establish effective anonymization. ICO guidance on effective anonymization.
These terms describe different outcomes or techniques.
| Term | What it means in a claims-data setting |
|---|---|
| Anonymization | Producing information from which people are no longer identifiable in the relevant context |
| Pseudonymization | Replacing or separating identifiers while retaining additional information that can reconnect records to people |
| Masking | Hiding or changing values, either in a stored dataset or in what a user can see |
| Redaction | Removing information from the copy being disclosed, such as deleting a name from a document |
| Tokenization | Substituting a reference value for an identifier, often with a separately protected mapping |
Replacing a claimant’s name with PERSON-1042 does not automatically make the record anonymous. If the insurer holds a table linking that reference to the person, the information remains personal data in its hands. Tokenization can support this kind of pseudonymization, but the mapping and the rest of the dataset need protection. ICO guidance on pseudonymization.
Masking can also be limited to a screen. A user may see only the last four digits of an account number while the complete value remains in the database. That controls visibility; it does not anonymize the underlying record.
The distinction matters because anonymization reduces the personal data an organization holds, while pseudonymization reduces the risks associated with identifiable data it still processes. ICO introduction to anonymization.
An insurer needs to examine both structured fields and the material attached to a claim.
Direct identifiers include names, email addresses, telephone numbers, government identifiers, and other values that point directly to a person.
Indirect identifiers may reveal someone when combined. An exact loss date, a small location, an unusual injury, and a distinctive occupation could describe only one person. Identifiability can arise from combining information with public records or other available data. ICO guidance on indirect identification.
In a claims file, the review may need to cover:
The scope includes other people mentioned in the claim, such as witnesses and injured third parties, as well as the policyholder.
The appropriate transformations depend on what the resulting dataset needs to support.
No technique is sufficient in every case. NIST’s de-identification guidance describes the need to combine suitable transformations with risk assessment and governance, and notes that tools that merely mask information may not provide adequate de-identification. NIST SP 800-188.
Suppose a fictional insurer wants to study how long household water-damage claims take to settle. Its operational records contain names, street addresses, exact dates, payment details, and adjuster notes.
A proposed analytical extract could use the following structure:
| Operational information | Possible analytical treatment |
|---|---|
| Claimant name and contact details | Omit |
| Street address | Replace with a broader region |
| Exact loss and settlement dates | Use reporting quarter and elapsed handling days where sufficient |
| Payment account details | Omit |
| Detailed narrative | Use a reviewed category such as “escape of water” |
| Claim amount | Retain, band, or aggregate according to the analysis and disclosure risk |
These are example transformations, not proof that the extract is anonymous. A very large claim in a sparsely populated region could remain recognizable. The team may need to combine categories, suppress unusual records, or release only grouped results.
There is also a tradeoff in usefulness. Converting all claim amounts into broad bands may prevent a severity analysis from measuring large losses accurately. The data design needs to preserve the information required for the actual question.
Claims-system testing often needs realistic relationships: a policy with several claims, a claim with multiple payments, or a reopened claim with revised reserves. It does not usually require real customer identities.
Purpose-built synthetic records can cover these scenarios. For example, a test suite can include a payment reversal, an expired policy, and two similarly named fictional claimants without copying a live claims file.
Synthetic data generated from real records still needs assessment for disclosure risk. It should not be assumed safe solely because a tool calls it synthetic. The ICO discusses synthetic data among anonymization techniques and the associated considerations. ICO guidance on anonymization techniques.
Where testing uses transformed production records, the process needs to cover attachments and exports as well as database tables. Otherwise, a masked record can still expose the original name through a linked document.
The relevant rules depend on jurisdiction, the organization, and the data involved.
In the United States, HIPAA provides two methods for de-identifying protected health information within its scope: Expert Determination and Safe Harbor. Safe Harbor includes removal of specified identifiers and a condition concerning actual knowledge of remaining identifiability. Expert Determination requires an appropriately qualified expert to assess and document a very small identification risk. These are not universal rules for every property and casualty claims dataset. HHS guidance on HIPAA de-identification.
Anonymization for analysis also does not mean deleting information from the operational claim file indiscriminately. The insurer may still need identifiable records to handle the claim and meet its retention obligations. The intended analytical output and the operational record serve different purposes.
If the insurer can reconnect an extract to individuals, that fact belongs in the assessment of whether the extract remains personal data. Calling a file “anonymous” or restricting access to it does not settle the question.
An effective review checks privacy and usefulness together. The privacy review asks whether people remain identifiable, including through rare combinations and linked information. The analytical review asks whether the transformed data still supports its intended purpose.
For the water-damage example, the team could compare the resulting claim counts and handling-time distributions with the source data. If excluding unusual records removes most long-running claims, the extract may give a misleading picture of settlement performance.
The transformation rules and the assessment should be documented. Changes to the recipient, data fields, or available linking information may require a fresh review. That gives the insurer a reasoned basis for using the resulting dataset and a clear explanation of what the data can support.

How the combined ratio measures underwriting performance, what its loss and expense components include, and why calendar and accident year figures differ.

What a core insurance system covers, how policy administration, billing and claims fit together, and where the boundary with surrounding tools sits.

What the insurance expense ratio measures, how it is calculated, which costs sit inside it, and why a falling ratio does not always mean lower costs.
Book a live, personalized demo with our product team - tell us your use case and see the platform work with your data. No commitment.