New Tax Declarations and Proofs – Now Fully Inside HR Blizz for India Payroll Read the release note
HR Blizz
All resources
March 24, 2026 · 11 min read
Global Payroll

What AI should do in payroll, and what it must never calculate

AI belongs in payroll as a check on the run, not the engine that computes the tax. Where the line sits, what the EU AI Act requires, and what to ask vendors.

Four vendor demos in seven weeks. Same deck order every time. Dashboard, a chart with a red spike on it, the letters AI in the corner of every slide.

There is one question worth asking in every one of them. Does your model calculate the tax, or does it check the tax after something else calculated it? The answer you hear often enough to expect it is that the model learns the country’s tax logic from historical payroll data, offered as the strong part of the pitch rather than the problem it is.

That question is the whole evaluation. If a model computes the withholding, the conversation is over. AI has a job in payroll, and that is not the job.

The German finance ministry publishes its algorithm every November

On 12 November 2025 the Bundesministerium der Finanzen published the Programmablaufpläne zur Lohnsteuer for 2026: the official program flow charts for machine calculation of German wage tax. Not guidance. The algorithm, step by step, that payroll software must reproduce. If your engine and the ministry’s own calculator disagree by one cent, your engine is wrong.

Three years from now an auditor will ask why employee 48812’s October 2026 Lohnsteuer was 743.18 euros. The answer has to be a chain you can walk: this taxable gross, this tax class, this allowance, this version of the rule effective on that date, that arithmetic. Rerun it, get 743.18 again. Not approximately. Again.

A probabilistic model cannot promise that. It produces a likely output given what it has seen, and retraining can move the answer, with no derivation you could put in front of a Finanzamt officer or, harder, in front of the employee whose rent depends on it. “The model is 96 percent confident” is not an answer to “why was I taxed this much.”

So be blunt. A vendor who says a model computes the tax is either using AI loosely for a rules engine, which is a marketing problem, or means it literally, which is a liability. Gross-to-net stays deterministic: coded rules, versioned by effective date, tested against the authority’s published cases.

What the model is genuinely good at is remembering last month

Payroll almost never fails at arithmetic. It fails at input, and at the gap between what the input said and what anybody expected.

A location code changes in the HR system and 40 people move tax jurisdiction. A benefits file arrives twice because someone re-ran the export after lunch. A recurring allowance stops appearing because the element expired on 31 December and nobody was watching. The engine calculated all of it perfectly. It calculated the wrong thing perfectly.

That is work for a model, not a person with a spreadsheet and four hours. Compare this period against prior periods, at element level, per employee, and surface what moved without a reason. A component present for eleven months and absent in the twelfth. A duplicate. A figure well outside its own trend. Nobody eyeballs 8,000 employees across 14 countries.

The second job runs after calculation: read the results back against that country’s statutory rules and flag what looks wrong before money moves. Contributions that ignored a ceiling. A net above gross. A minimum wage floor breached because hours were coded unpaid. The rules decide; the model notices and ranks. Two smaller jobs pay too: parsing messy inbound files, so a timesheet extract with renamed columns gets mapped rather than rejected at 11 p.m., and answering pay questions from documented rules, with a citation.

HR Blizz splits it that way. Gross-to-net calculates natively per country and is deterministic; the AI runs twice per period around it. Before calculation it compares the run against the last 12 periods for duplicates, vanished or new elements, and figures more than 30 percent off their own history. After calculation it reads each payslip against that country’s statutory rules and names the employee, the reason and the fix. Both passes see an employee ID and numbers, never a name.

The bank file is a one-way door

EY’s December 2022 survey of 508 US payroll respondents found that one in five US payrolls contains errors, and that the average organisation makes 15 corrections per pay period. Those are errors caught inside the process. Catching one after the payment file has cleared is a different problem, and the difference is legal rather than operational.

Recovering an overpayment is governed by employment law in each country, and the rules are neither friendly nor consistent. In the UK, section 14 of the Employment Rights Act 1996 carves overpayment recovery out of the unlawful deduction rules, which sounds generous until the employee has spent the money and you are arguing about change of position. Elsewhere you need documented consent, or a repayment schedule capped so it cannot push net pay below a protected floor, or you are out of luck for the portion already remitted to a social fund.

An error caught on a Wednesday is a corrected line in a run. The same error caught the following Monday is a legal project.

Annex III does not say “payroll”. That is the part worth reading carefully.

Annex III point 4 of the EU AI Act covers employment and worker management. Point 4(a) captures recruitment and selection. Point 4(b) captures systems used “to make decisions affecting terms of work-related relationships, the promotion or termination of work-related contractual relationships, to allocate tasks based on individual behaviour or personal traits … or to monitor and evaluate the performance and behaviour of persons in such relationships.”

Read that against a payroll variance check. A model comparing this period’s overtime line to the last twelve and raising a flag for a human is not deciding terms of work and is not evaluating anyone’s performance. Article 6(3) is explicit: an Annex III system is not treated as high-risk where it detects deviations from prior decision-making patterns and is not meant to replace or influence a completed human assessment without proper human review. Profiling of natural persons is high-risk regardless, with no exception. Scoping is not an escape route: under Article 49 a provider who concludes its system is not high-risk still registers itself and that assessment in the EU database. Scoping sets the size of the obligation and forces both sides to write down what the system may do.

Where high-risk applies, the obligations are concrete. Under Articles 9 to 15 the provider carries risk management, data governance over training and test data, technical documentation, automatic logging, instructions for use, oversight designed into the product, and accuracy, resilience and cybersecurity, plus conformity assessment and registration. You carry Article 26: use the system per those instructions, assign oversight to named people with the competence, training, authority and support to exercise it, keep your input data relevant and representative, retain logs at least six months, report serious incidents, and inform workers’ representatives and affected workers before the system goes to work. Article 99 sets fines up to 15 million euros or 3 percent of worldwide turnover.

On timing, be careful, because this is moving underneath you. As the law stands today, Annex III obligations apply from 2 August 2026. The Commission’s Digital Omnibus on AI, published 19 November 2025, proposes deferring standalone high-risk obligations to 2 December 2027. It is a proposal in negotiation, not law: the Council agreed its position on 13 March 2026, Parliament’s IMCO and LIBE committees adopted a joint report on 18 March, and the plenary mandate and trilogues come next. Nothing is deferred until they land. The Commission also blew its own 2 February 2026 deadline for the Article 6(5) guidelines on classifying high-risk systems, and seven weeks later there is still no draft. Treat any classification you hold as provisional, and check where the Omnibus stands before you rely on a date.

US rules are a patchwork aimed at hiring, not paying: New York City’s Local Law 144 bias audit, Illinois HB 3773 since 1 January 2026, Colorado’s SB 24-205 pushed to 30 June 2026. None bite on a variance check no employment decision rests on. All bite the moment a platform recommends who gets what.

What the approver does after the machine has already looked

Approval used to mean recomputing enough of the run by hand to believe it. Now it means adjudicating a list someone else produced. Article 14 names the risk: automation bias, over-relying on output from an automated system, especially one framed as a recommendation. Oversight is not a person clicking approve. It needs someone who understands the system’s limits, can interpret its output, and can override or stop it.

Which makes the flag count the number to interrogate. A check that returns 400 items on an 8,000-employee run is not thorough, it is broken. Nobody triages 400 items on the Tuesday before pay date. They spot-check twelve, approve, and the one real error sits in the other 388. Financial crime compliance learned this expensively: FATF and industry data put false positive rates in AML transaction monitoring as high as 95 percent, which is why banks staff floors of analysts closing alerts that were nothing.

Sixteen flags, each naming the employee ID, the rule, expected versus actual and a recommended action, gets worked properly. Ask for the alert count on a run of your size and the share of last quarter’s alerts that were real. If they cannot tell you, they are not measuring precision, which means they optimised for looking thorough.

Three data questions decide whether legal signs off. Get them answered in writing. Does our pay data train a model shared with other customers, or does the model read only inside our tenant. Does any of it leave the region we contracted for. Does the model see names, or identifiers and amounts. An anomaly check does not need to know that employee 48812 is Klara, and what the model never sees is risk you never have to manage.

Four questions for the next demo

Ask these.

1. Does a model compute any statutory amount on a payslip? If not, show me the rule source and version history for one country’s income tax.

2. What does the AI do, where in the cycle, and does it run before the payment file is created or after?

3. On a run our size, how many items does it flag, and what share of last quarter’s flags were real?

4. Which parts of your product are high-risk under Annex III, what is your Article 6 assessment, and what does your instructions-for-use document require of us?

Then do one thing internally. List every error from your last three periods that reached payment, with what it cost to unwind, and sort by cost. That list, not the vendor’s deck, tells you what a check has to catch.

HR Blizz keeps gross-to-net in coded statutory rules on its own engine, with AI checking each run twice on identifier-only data before money moves. If that is the shape you want, talk to our team.

A note on the regulatory detail above: it is general information rather than legal advice, and the AI Act timetable in particular is still moving.

FAQ

Q: Can AI calculate payroll taxes?

It should not. A statutory gross-to-net calculation has exactly one correct answer for a given set of inputs, it must be reproducible years later in an audit, and it must be explainable to both the employee and the tax authority. Germany’s finance ministry publishes the wage tax calculation algorithm each year for software to reproduce exactly, which is the clearest illustration of why the computation belongs in deterministic coded rules rather than a probabilistic model.

Q: What does AI actually do well in payroll?

Four things: comparing the current run against its own history to catch duplicates, disappeared pay elements and figures far off trend; reading calculated results against statutory rules to flag likely errors before payment; classifying and parsing messy inbound files; and answering employee pay queries from documented rules. All four are detection and support tasks with a human deciding, not calculation tasks.

Q: Is payroll AI high-risk under the EU AI Act?

It depends on function, not on the label. Annex III point 4 covers AI used in recruitment and in decisions affecting terms of employment, promotion, termination, task allocation and performance monitoring. Article 6(3) provides that a system intended to detect deviations from prior patterns, without replacing or influencing a completed human assessment, is treated differently, though anything performing profiling of natural persons is always high-risk and any such assessment still has to be documented and registered under Article 49.

Q: When do the EU AI Act high-risk obligations apply?

As the Act currently stands, obligations for Annex III high-risk systems apply from 2 August 2026. The Commission’s Digital Omnibus on AI, published on 19 November 2025, proposes deferring standalone high-risk obligations to 2 December 2027, but as of late March 2026 it was still in negotiation, with the Council position agreed on 13 March 2026 and Parliament’s mandate and trilogues still ahead. It should not be relied on until adopted.