CASE STUDY / 02
AI voicebot for telecom customer care
Backend development, reliability and database optimisation for customer care at national scale.
- Role
- Backend developer · Database administrator
- Context
- Telecommunications · Customer care
- Period
- 2021 — 2024
A professional experience described through my own contribution. Client code and documentation are not published on this page.
My contribution at a glance
- Backend developer and database administrator on an IBM watsonx-based AI voicebot for customer care at one of Italy’s largest telecom providers
- Led a 3-month data migration from the operational (flow) database to a federated Analytics database, building PL/SQL routines to move and remodel the data for compatibility and performance in downstream visualization dashboards
- Detected and resolved a critical memory leak before it caused a production outage; introduced Resilience4j circuit breakers, cutting platform error rate from ~12% to under 3%
- Led database refactoring during cloud migration (indexing on 500M+ record audit tables, connection pooling, partitioning), improving query execution time by 85% and reducing DB CPU usage by 60%
The night a memory leak almost took down customer care.
Backend developer and database administrator on an IBM watsonx-based AI voicebot for customer care at one of Italy’s largest telecom providers, over three and a half years.
The customer at the centre of the service
A customer care voicebot is the first point of contact between a customer and the telecom operator: it takes the request, handles it in conversation and relies on business systems to give a concrete answer. That is why every technical decision had to be read from the caller’s side: an error or a long wait means an unresolved request and a customer who has to repeat themselves or be passed to an agent. I worked on the part the customer never sees but that shapes their experience: the backend, behaviour under load, audit tables with over 500 million records and the Analytics database feeding the service dashboards.
A problem I caught before it became an outage
During routine monitoring I spotted a pattern that shouldn’t have been there: rising RAM usage, CPU spikes at peak traffic, slower and slower query responses. The cause was a module opening database connections without closing them properly — heading toward a full outage on a system that customer-facing voice interactions depended on. I coordinated a controlled shutdown of the affected service, fixed the connection lifecycle, and added idle-timeout thresholds — without taking the platform down. Afterwards I pushed for continuous connection-pool monitoring to become standard practice, and introduced circuit breakers with Resilience4j that cut the platform’s overall error rate from around 12% to under 3%.
Migrating to the cloud meant questioning old assumptions
When we moved from on-premise to IBM Cloud, inefficiencies in the existing database access patterns surfaced that had been invisible before. Three options were on the table — lift-and-shift, schema refactoring, or a rewrite of the orchestration layer. I built the case for refactoring using actual query execution data rather than general recommendations: queries over 8 seconds were becoming systemic bottlenecks, and inefficient connections risked a significant infrastructure cost increase. The refactoring — targeted indexes on audit tables with over 500 million records, connection pooling, table partitioning, query rewrites — cut query execution time by 85% and database CPU usage by 60%.
A separate migration, a different challenge
Over three months, I led a data migration from the operational (flow) database to a federated Analytics database, building the PL/SQL routines that moved and remodeled the data so it would be both compatible with and efficient for the visualization dashboards downstream.