Blog Best practices

AI in Debt Collection: Trends, Benefits and Risks in 2026

AI in Debt Collection: Trends, Benefits and Risks in 2026

This guide covers what Artificial Intelligence in debt collection really does, how the scoring layer decides who to contact first, where the technology helps and where it doesn't, and three modern solutions worth knowing: HES CollectionAgent, Symend, and InDebted.

U.S. household debt reached $18.8 trillion in the first quarter of 2026, and about 5.0% of consumers now carry a third-party collection account on their credit report, according to the Federal Reserve Bank of New York. Volume is rarely what defeats a recovery team.

The harder question is which of tens of thousands of delinquent accounts to work today, on which channel, and with what message, without spending agent hours on people who were going to pay anyway or leaning on people who genuinely can't. That is the problem Artificial Intelligence in debt recovery is built around. The advance that matters is not faster dialing. It is a sharper score.

What Artificial Intelligence in Debt Collection Actually Means

AI in debt collection is a set of techniques applied to recovery, not one product. In practice it bundles four things that get blurred together in marketing: machine-learning models that estimate repayment probability, natural language processing that reads and drafts messages, predictive analytics that rank and segment accounts, and newer agentic systems that pick the next action with limited human input.

Why separate them? Because they carry very different risk. A model that ranks accounts by likelihood to pay is well understood, and a competent team can audit it. A conversational agent negotiating a payment plan on a live call is another matter, with compliance and reputational exposure a ranking model never touches. Anyone evaluating AI-powered debt recovery software should pin down which of the four a tool actually performs, and which it only puts on a slide. Before any of that matters, it is worth being precise about what the current process costs.

Where Traditional Debt Collection Breaks Down, and What AI in Debt Collection Fixes

Most collection operations still build the daily work queue from two fields: how far past due an account is, and how much it carries. That ordering is simple to run and simple to defend in a portfolio review, and it says nothing about whether the borrower can pay or intends to. The cost lands in two places. Collector time goes to accounts that were never going to clear, while accounts that would have resolved after one properly timed contact sit untouched until they age into a harder tier.

The consequences are documented in the federal complaint record. The CFPB received roughly 387,400 debt collection complaints in 2025, and the most frequent allegation was an attempt to collect a debt the consumer states is not owed. That category has ranked first in every year since the Bureau began accepting debt collection complaints in 2013. A pattern that stable points less to collector conduct than to the records behind it: balances never updated after a payment, accounts settled but still active in the queue, and the wrong person tied to an open file. Where those records drive outreach, a measurable share of collection spend produces disputes instead of recoveries.

The economics compounds the problem. Collections carries a largely fixed payroll against a portfolio whose recoverability erodes the longer an account stays open, so the same headcount returns less as the book gets older. Consumer lending also lacks a common public benchmark for cost-to-collect. Definitions vary between firms and few disclose the figure at all, which leaves a CFO comparing the ratio against nothing except prior quarters of the same operation.

Four bottlenecks account for most of it:

  • One strategy for everyone. The same script, cadence, and channel reach a borrower who missed one payment after a job change and one who has ignored four cycles. Both get treated as the portfolio average, and neither behaves like it.
  • Collector hours pointed the wrong way. Manual queues favor large balances and old cases, which recover worst, while dead numbers and stale addresses burn attempts before any conversation starts.
  • Behavioral data that never reaches a decision. Lenders hold years of payment timing, channel response, and broken-promise history. Without a model reading it, that record sits in a warehouse while the queue is built from two fields.
  • Controls that depend on someone remembering. When call timing rules and contact caps live in training material instead of the workflow, every exception becomes an individual judgment call, and those exceptions are what examiners find later.
Operating questionTraditional collectionsAI-led collections
What sets queue order?Days past due and balanceExpected recovery per account
How is contact data treated?Trusted until an attempt failsScored for reachability first
When is hardship identified?When the borrower says soBefore escalation, from behavior
What triggers escalation?Fixed calendar stepsA change in repayment behavior
When does the team see results?In the next reporting cycleWhile the strategy is still running

What these failures share is a problem of sequencing rather than effort. Which accounts to work on, and on what evidence, has to be settled before anything about channel, script, or timing, and it is the decision manual queues get wrong most consistently. Closing that gap is what AI in debt collection is for, which makes it worth being precise about what the technology does and does not do.

How AI-powered Debt Collection Works: The Three Layers

Most working AI debt collection systems are built the same way, on three layers that run in sequence: the data a model reads, the score it produces, and the action that score sets off.

Layer 1: The Data That Predicts Repayment

The model reads account-level data: payment and delinquency history, transaction patterns, prior contact and response records, and, where regulation permits, enriched behavioral or third-party signals. Breadth is the point. A days-past-due bucket tells you almost nothing about whether someone pays next week, whereas a wider feature set begins to separate the temporarily stuck from the genuinely distressed.

Layer 2: Scoring and Prioritization

Here the data becomes a score, usually a propensity-to-pay or collectability estimate, and the portfolio re-ranks as fresh information lands. This layer holds the economic value of any AI debt collection system, so it gets its own section below.

Layer 3: Adaptive Action

The score then triggers a next step: a reminder, a channel switch, a payment-plan offer, an escalation, or no contact at all. Agentic tools keep adjusting the sequence as behavior shifts instead of marching down a fixed dunning calendar. Whether an AI debt recovery tool advises a collector or acts on its own comes down to this layer, and so does the need for tight governance.

AI Tools Used in Debt Collection, Layer by Layer

Vendors can present these three layers under a longer list of names. We mapped them to the decision each one actually makes, together with the question worth putting to a vendor at the demo.

ToolWhat it decidesWhat to watch for
Propensity and outcome modelsWho repays, relapses, or rolls forwardTrained on past contact, so old bias carries forward
Dynamic re-scoringWhen an account changes priority mid-cycleWorthless if event data arrives on a nightly batch
Next-best-action enginesChannel, timing, and cadence per accountFrequency caps must bind, not advise
Conversational and voice AIRoutine inbound queries, without a collectorTCPA consent applies to synthetic voice
Generative AIWhich approved wording fits this borrowerOnly inside human-approved language
Agentic systemsThe next step, with limited human inputEvery action needs a logged reason
Causal and uplift analyticsWhich action actually changed the outcomeNeeds holdout groups, which teams skip

The Scoring Engine Behind AI/ML Debt Recovery: Who Pays, and Who Won't

Coverage of AI in debt collection tends to concentrate on the conversational layer. The score determines which accounts are worked and in what sequence, while everything downstream of it, the calls, messages, and payment links, executes a decision that has already been taken. A capable voice agent operating on a weak score does not correct that decision; it carries it out more efficiently.

From Days-Past-Due to Propensity to Pay

Traditional recovery sorts accounts by age and balance: 30 days, 60 days, 90 days, biggest exposure first. Easy to run, easy to defend in a meeting, and blind to the thing that actually predicts recovery, which is whether a given person can and will pay.

A propensity-to-pay model, the heart of any AI debt collection system, estimates that probability from payment history, behavioral signals, and past outcomes, then ranks the book by expected recovery rather than by calendar position. Mature setups go past a single number and score several outcomes at once, such as likelihood to repay, to relapse after a promise, or to respond on a given channel, and they re-score continuously as new events arrive.

McKinsey has reported recovery-rate improvements of 10 to 15% and collections-efficiency gains of 30 to 40% from analytics of this kind, alongside a North American bank that saved roughly $25 million on a $1 billion portfolio after deploying machine-learning self-cure models.

Separating "Can't Pay" From "Won't Pay"

The single most useful move an AI-powered debt collection model makes is to split two groups that look identical on a days-past-due report. One can't pay right now: hardship cases who need a plan or forbearance. The other can pay and chooses not to. Treat them the same way and you lose on both.

The right response runs in opposite directions, flexibility for the first group, firmer and earlier escalation for the second, and a model that scores capacity and intent separately lets a team route hardship away from aggressive outreach.

Where regulation allows, enriching internal records with behavioral and digital signals builds a fuller debtor profile and flags the accounts not worth chasing at all, such as likely fraud or numbers no one will ever answer.

Why Explainability Decides Whether the Score Is Usable

A score a credit committee can't explain is a score that won't survive an exam. A regulated lender needs the model to give reasons, not just a ranking: which factors moved a given account, whether any input quietly proxies for a protected class, and how each decision is logged for later review.

A model that improves recovery but can't be explained to an examiner tends to stall in model-risk review, however good its numbers look. That is the practical reason opaque scoring keeps losing ground to approaches that show their work, and it should carry real weight in any debt collection software decision.

Where Artificial Intelligence Reshapes the Debt Collection Workflow

With the score in place, the AI-powered collection brain changes how the rest of the work runs.

Personalized Outreach at Scale

Inside an AI debt collection platform, the next-best-action engine replaces a one-size dunning script, picking the channel, timing, and frequency most likely to land with each segment, and backing off when contact starts to fatigue.

McKinsey's customer research found that contact preferences track personal habit far more than the risk tier a lender assigns, which is why a fixed call-then-letter cadence loses to a model-chosen one. The gain is real but it has a ceiling: better sequencing lifts contact and response rates, it does not create ability to pay.

Conversational and Voice AI Agents

This is the noisiest corner of the "smart" debt collection market, and the one to question hardest. Conversational agents can field routine inbound queries, send a payment link, and update a record without pulling in a collector, which frees people for disputes and hardship negotiations.

The catch is that they operate inside a regulated conversation, so approved language, consent, and frequency caps have to be coded as hard limits the system can't quietly override. On consent, the FCC has already drawn the line: it confirmed in February 2024 that an AI-generated voice counts as an "artificial voice" under the TCPA, so a synthetic-voice call placed without prior express consent is a violation on its own, with statutory damages that run per call.

Compliance Enforced in Real Time

The strongest use of AI in collections right now might be compliance itself. A debt collection platform can check disclosures, calling windows, and contact frequency against FDCPA, Regulation F, and TCPA as a conversation happens, then log every action for the audit trail.

Oversight moves from after-the-fact sampling to continuous enforcement, which matters in a function where a single bad pattern can erase a meaningful share of what the team recovers.

What Artificial Intelligence in debt collection still can't fix

A guide that only lists upside is a brochure. Three limits of AI debt collection deserve a skeptic's attention.

The first is uncomfortable for the whole category. A Yale School of Management study of roughly 22 million collection cases found that borrowers contacted first by an AI caller repaid less and broke their repayment promises more often than those handled by people, and that later human follow-up never fully closed the gap.

The researchers tie part of it to the plain fact that the borrower knew they were talking to a machine. AI still earns its place in early contact. What the study warns against is leaning on it at the exact moment a borrower commits to pay. Keep a person on that conversation, then watch whether the promise actually holds and change tack the moment it slips.

Second, an AI debt collection model is only as good as its data, and collections data is often patchy, out of date, or skewed by who got contacted in the past. Train on historical outcomes without checking for that, and the model learns the bias along with the signal, which is an ethics problem and a regulatory one at the same time.

Third, automation without governance adds risk to a debt recovery rollout rather than removing it. An agent that contacts the wrong person at the wrong hour scales a breach as fast as it scales recovery. Teams that get value out of AI debt collection software build oversight, bias testing, and explainability into the system from day one.

What the Law Now Demands of AI-driven debt recovery: the US, EU, and UK in 2026

The compliance question has changed. A year ago the argument was whether to let a model near recovery at all. In 2026 the moving part is the rules that govern that model, and they move at a different pace in each market. None of the three major regimes bans the technology outright. Each one decides which uses draw extra scrutiny, and a lender that operates across borders needs to know where its system stands in all three.

United States: Old Statutes, New State Rules

The United States has no single AI statute for collections. What binds is existing consumer-protection law applied to new technology, with a fast-growing layer of state and city rules on top. The federal hooks are already on the books, including the TCPA consent rule for synthetic voices noted above and the duty under ECOA to explain the basis for an adverse decision.

The newer movement is local. As enforcement at the CFPB has pulled back, cities have stepped in, and New York City's SHIELD Rule, effective September 1, 2026, is the one to read first. It caps collectors at three contact attempts in seven days across calls, texts, and emails combined, lets a consumer dispute at any point, and stops collection until the debt is verified within 60 days. For an automated outreach engine that means tracking attempts across every channel at once and halting the moment a dispute arrives, both written in as rules the engine cannot bypass. Colorado has gone further with a law aimed at algorithmic discrimination in consequential decisions such as lending, and more states are drafting their own.

European Union: The One Statute Written for the Technology

The EU AI Act is the only regime built specifically for AI, and for recovery work the trigger is the score. A model that evaluates creditworthiness or sets a credit score is classed high-risk, though fraud-detection systems are carved out. That status brings duties on documentation, data quality, human oversight, and record-keeping.

The deadline most lenders planned for then shifted: under the Digital Omnibus agreed in May 2026, those obligations now point to December 2, 2027 rather than August 2026. Two neighboring rules stack on top. The obligations that landed on general-purpose model providers in August 2025 sit upstream of a lender but shape the documentation a deployer should demand, and under DORA the same model counts as an outsourced ICT service, so a single per-request log can satisfy both.

Two things did not move: the duty to disclose that a person is dealing with a machine still lands in August 2026, and the GDPR already bites, since the EU's 2023 Schufa ruling treats a determinative score as an automated decision the person can have explained and reviewed. A lender outside the EU with borrowers inside it is in scope regardless of base.

United Kingdom: Rules by Use, not by Label

The UK has no AI statute and, as of 2026, no plan to write one. The FCA supervises the technology through rules it already has. Its Consumer Duty asks any firm whose models shape pricing, eligibility, or the treatment of customers in financial difficulty to show the outcome was fair, with a named senior manager answerable for it.

On the data side, the Data (Use and Access) Act 2025, in force since February 2026, lets a solely automated decision that significantly affects someone stand only with safeguards: the person has to be told, be able to get a human to review it, and be able to challenge it. A decision to escalate an account or write a debtor off sits right inside that test.

Read side by side, the three regimes ask for the same short list, whatever the market: a decision the firm can explain, a person who can step in, a record of what the system did and why, consent before automated contact, and proof the model was checked for bias.

A program built to the strictest of the three usually clears the other two. That points to one practical move: keep the scoring layer transparent and the message content under human-approved control, rather than trust either to a system no one can read.

The Business Case for Artificial Intelligence in Debt Collection

Stripped of the hype, the logic behind AI debt collection is plain. AI points scarce collector hours at the accounts where those hours change the result, and hands the routine contact to software.

The direction is well supported: McKinsey's collections work points to double-digit recovery gains, 30 to 40% efficiency improvement, and machine-learning self-cure models that free 5 to 10% of collector capacity. How much of that a given lender captures depends on portfolio mix, data quality, and how disciplined the rollout is, and it shows up sooner in early-stage delinquency than in late-stage or charged-off books.

For a board, the credible promise is a measurable lift in recovery and a lower cost-to-collect inside a defined payback window. Anyone pitching it as a wholesale reinvention of collections is overselling. A sober first step is to run an AI debt collection model in parallel with the current process on one portfolio segment and compare recovery and cost head to head before scaling the program.

Modern AI-powered debt collection software, by best-fit scenario

The AI debt collection software market sorts by what a tool is built around, not by who wins overall. The three below occupy different scenarios. What follows is independent desk research from publicly available materials, current as of June 2026, with the same questions put to each vendor; confirm current details directly before deciding.

HES CollectionAgent: full-lifecycle scoring and decisioning

HES CollectionAgent is built around the scoring layer this guide centers on, rather than treating collections as a module bolted onto a lending suite. Its machine-learning engine scores debtors on repayment probability once an account turns delinquent, separates "can't pay" hardship cases from "won't pay" strategic defaulters, and flags accounts not worth pursuing such as likely fraud or unreachable profiles. Where regulation allows, it enriches internal data with behavioral and digital signals for a fuller debtor profile.

On top of the score sits a next-best-action engine that chooses channel, cadence, and timing across email, SMS, push, and voice, and re-routes in real time when behavior changes.

Its promise-to-pay monitoring is the part that speaks directly to the Yale finding above: it tracks each commitment, and when a payment fails it shifts the account into a softer-pressure or retention track without waiting for a human. Accounts are categorized automatically by days past due, with portfolio and account-level dashboards for visibility.

Pre-approved message templates and configurable no-code workflows let a team change rules and strategies in minutes, which is the part CTOs tend to care about. It runs API-first within an ISO/IEC 27001 framework, validates every automated action against rules such as GDPR and FDCPA, and deploys on AWS, Google Cloud, or on-premises.

HES reports response-rate gains near 50%, recovery improvements around 25%, and up to 90% lower operational cost and processing time; treat those as vendor-reported and test them on your own book.

Best for: regulated lenders and BNPL providers that want scoring, decisioning, and outreach as one controllable AI debt collection system. Watch for: lenders that also need full origination-to-servicing should check how it sits next to a broader loan platform.

Sources: HES FinTech product and launch materials, accurate as of June 2026.

Symend: behavioral-science engagement

Symend's platform, SymendCure, is built around behavioral science for early-stage delinquency. It scores accounts by likelihood to repay, sorts them into "delinquency archetypes," and runs digital engagement journeys across email, SMS, IVR, and self-serve portals using tactics such as loss-aversion framing and social proof. It works as a digital engagement layer rather than a voice-calling or full-lifecycle platform, and its strongest published case studies sit in telecom and utilities. Symend reports up to 10% higher recovery rates and sizable cost cuts; these are vendor-reported.

Best for: first-party creditors focused on early-stage, relationship-preserving digital engagement. Watch for: it is not a voice-AI or late-stage recovery tool, and telecom results may not transfer to other verticals.

Sources: Symend public materials and an independent third-party review, accurate as of June 2026.

InDebted: consumer-centric digital recovery

InDebted's product, Collect, is built around the consumer experience in third-party recovery, with machine-learning models tuning outreach across SMS, email, and chat, plus an AI Collector that uses conversational AI for inbound queries. InDebted reports that a large share of inbound requests in some markets resolve through conversational AI, and that its AI-written messages raised conversion; these figures are vendor-reported and partly market-specific. A human customer-experience team handles escalations and vulnerable customers.

Best for: lenders outsourcing consumer recovery who want a digital-first, experience-led approach with a human off-ramp. Watch for: it leans toward an outsourced recovery service rather than software you run in-house; confirm deployment options.

Sources: InDebted public materials, accurate as of June 2026.

ESTIMATE YOUR LENDING PRODUCT IN 16 STEPS
STEP: /

The Future of AI in Debt Collection: What Changes by 2027

Three shifts are already visible in procurement conversations, and together they describe what a defensible program will look like by the end of 2027.

The first is autonomy. Agentic systems today mostly recommend, with a collector approving. As they begin to act unsupervised, the open question stops being capability and becomes accountability: who owns a decision that no person reviewed, and what record shows it was reasonable. Firms that settle that in their own operating model, rather than leaving it to a vendor contract, will move faster, because the answer is what determines how much autonomy a risk committee will actually approve.

The second is a change in what gets modeled at all. Propensity scoring has been the standard for a decade and answers who will pay. Treatment optimization answers the harder question of which action changed the result, and it demands the holdout discipline described above. Adoption will run slower than vendors suggest. It is also where the remaining efficiency sits once prioritization is solved.

The third is a date. December 2, 2027 is when EU high-risk obligations bite for creditworthiness scoring, and because the evidence they demand closely resembles what US examiners and the FCA already look for, that standard is becoming the practical global bar for what a lender must be able to show. Building to it now avoids a rebuild later.

FAQ

Is the use of Artificial Intelligence in debt collection compliant with FDCPA and Regulation F?

Does AI replace human collectors?

How does the scoring work in debt collection?

What are the main challenges of traditional debt collection?

What AI tools do collections teams actually use?