Axoly Tech

From audio to text: AI, privacy, and asynchronous processing.

In the digital age, data privacy has become an essential issue. That's why we build ethical, sustainable digital services designed to protect your data and to last over time.

The project

Recently, we took on an exciting challenge: designing an AI-powered solution that preserves technological sovereignty and users' privacy.

Rather than transferring sensitive audio files to external servers, we propose a decentralized architecture that lets each organization deploy the tool directly on its existing infrastructure. This is a transcription service that offers an alternative to traditional software.

Illustration of the “From speech to text” service: a form for choosing an audio file and entering an email address, surrounded by people recording a podcast and a person reviewing their transcription.

Strategic objectives

1

Guarantee maximum data privacy: whether for sensitive data, a need for confidentiality, or simply because we don't want our recordings ending up on third-party platforms. This calls for greater autonomy.

2

Promote an eco-design approach: reusing existing computer hardware avoids overconsumption of resources and planned obsolescence.

The principles behind our solution

Accessible to everyone

Every member of the organization can easily convert an audio file into written text.

Open to multiple people at once

The software must allow the service to be used simultaneously, without limitation or constraint.

No processing urgency

An instant transcription isn't strictly necessary. Users can submit their files and retrieve the transcriptions later.

Integrated at no extra cost

The solution must integrate seamlessly with the existing infrastructure, without generating additional costs after its initial development.

The software

  1. Each employee logs into the application's website and uploads their audio file, along with their email address to receive the transcription.

  2. The file is saved on the server.

  3. At night, when the company's server is less busy, a batch process kicks off. A transcription is generated for each file uploaded during the day, using a specific AI module.

  4. Once the transcription is complete, the server sends an email containing the transcription to the person concerned.

Diagram of how the software works: the audio file is sent to the server, transcribed asynchronously overnight, then the completed transcription is emailed to the person concerned.
Diagram of the application's components: the web application uploads audio files processed by the backend, a batch process retrieves them from the database, transcribes them via the audioToText AI model, then the email-sending component delivers the transcription to the person concerned.
Diagram of the application's components

Methodology

To develop this application, we used an agile methodology and incorporated the theoretical framework of the GREENSOFT model. A Sustainability Journal was created, in which we recorded every decision and its consequences in terms of sustainability.

This approach makes it possible to:

  • integrate sustainability concerns and requirements into the development process;
  • keep a record of the changes and implementations made, along with their impacts in terms of eco-design.

Eco-design

We can't ignore the environmental impact of AI models: both their training and their execution consume a lot of energy and hardware resources. That's why we placed digital sobriety at the heart of our approach. Every technical decision, from server choice to model choice, is evaluated in order to design a genuinely useful solution while limiting its footprint.

How do we achieve our sustainability goals?

  1. Prioritizing efficiency over instant responses

    • Using or reusing less powerful servers: transcription takes longer, but there's no need to create new servers or incur additional costs.
    • Limiting the impact on hardware resource consumption. Most software runs during peak energy demand. We decided to run it at night, which optimizes server usage.
    • Decoupling the sending of files to the server from their processing.
  2. Determining the AI model

    • Autonomy and independence: open AI models that can be replaced.
    • Pre-trained models specialized in specific languages: this lets us achieve better transcription, even with smaller models.
    • Following testing, we identified the lightest model capable of producing acceptable transcriptions.

What's left to do?

Feature level

  • Expand the range of potential AI models: text translation or summarization of large documents.
  • An independent module adaptable to any type of task requiring asynchronous processing.

Eco-design level

  • Quantify the impacts: measure the carbon emissions generated per minute of transcribed audio.
  • Meet the essential criteria of GR491. As this is a pilot project, only a few criteria were considered during the development phase. A more thorough analysis would be needed if the application were to evolve further.

The protection of your data and sovereignty

Through the development of this software, we demonstrated that it's possible to deploy services using AI models on standard servers, without relying on external services, thereby preserving data privacy and sovereignty. And this, without incurring additional costs once the system is implemented.

What's more, by questioning the actual use of technological tools, we can observe direct impacts linked to their implementation. In our case, by choosing to process our transcriptions asynchronously and with a delay, we reuse existing servers that would not have been sufficient had an immediate response been required.

This approach illustrates how a thorough reflection on deployment methods can lead to solutions that are more efficient, resource-sparing, and respectful of existing resources.

Your project has a history. Let's build what comes next, together.

A first 30-minute conversation, with no commitment, to see whether we can help.