← Research areas

Research in progress

Research on AI processors with persistent NAND memory and local computation

The project studies architectures in which AI model weights are stored in non-volatile NAND memory and used by local computation circuits. Its research perspective includes a NAND–CMOS chiplet architecture and integration with 3D NAND memory.

Research objective

Explore a self-contained system capable of running AI models with approximately 7 billion parameters, with a target speed of 20 tokens per second and a target power of 30 watts. The research must establish whether these targets can be reached and under what conditions.

Weights are the numerical values that describe a model’s learned behaviour. A token is a unit of text processed or generated; tokens per second indicate generation speed.

Approach

The studied data path connects persistent weight storage to data transfer and computation. The double buffer provides two temporary memory areas to organise alternating data availability; a controlled MAC unit performs multiplication and accumulation operations.

Data path

  1. NANDPersistent weights · INT4
  2. Double bufferData transfer and alternation
  3. MACMultiply and accumulate · INT8 activations
  4. ResultsOutput values · INT32
Functional diagram of the demonstrator. MAC computation takes place in the FPGA, a device with reconfigurable logic.

In the demonstrator, computation takes place in the FPGA. The NAND–CMOS chiplet architecture and integration with 3D NAND are research perspectives separate from this result.

Preliminary results

Results obtained in simulation

The demonstrator has been verified in simulation on four 16×256 matrices, using INT4 weights, INT8 activations and INT32 results. These names refer to integer representations with 4, 8 and 32 bits, respectively.

Data-path tests
150
Successfully completed in simulation.
Tests with non-uniform matrices
16
Successfully completed in simulation.
Conditions of the CMOS MAC model
27
Successfully verified in the simulated model.

These checks concern the demonstrator and the CMOS MAC model. They do not yet demonstrate execution of the approximately 7-billion-parameter model or the complete processor’s targets of 20 tokens/s and 30 W.

Research status

Simulation, FPGA logic synthesis and experimental validation are distinct levels of verification. Results at one level do not replace the evidence required at the others.

Checks completed

Simulations

Preliminary results are available for the demonstrator on four matrices and for the CMOS MAC model conditions described on this page.

Findings to be documented

FPGA logic synthesis

Synthesis translates the logic description into an implementation intended for the FPGA. Resources, constraints and timing reports must be documented separately; they are not experimental measurements on the prototype.

Still to be performed

Experimental validation

Building the prototype and taking hardware measurements are the next required checks. The complete processor’s targets have not yet been demonstrated.

Next activities

Moving from simulation to a prototype will allow operation to be checked under actual conditions and measurements to be collected for comparison with the research targets.

  1. Prototype implementation

    Build the hardware demonstrator and prepare experimental tests of the NAND, double-buffer, MAC and results data path.

  2. Timing and correctness

    Verify transfer and computation timing and measure result correctness on the prototype.

  3. Bandwidth, current and energy

    Measure data bandwidth, current draw and energy, keeping experimental results separate from design estimates.

  4. Interface and industrial feasibility

    Study the internal NAND–CMOS interface and assess the industrial feasibility of the chiplet architecture and integration with 3D NAND.

Documentation

This section is reserved for research documentation. The technical dossier, diagrams and methodology will be linked here when available for consultation. No downloadable files are currently published.

Technical dossier

Description of the demonstrator, preliminary results, architectural hypotheses and verification limits.

Diagrams

Data path, double-buffer organisation and the relationship between memory and computation units.

Methodology

Matrix configuration, numerical formats, data-path tests and CMOS MAC model conditions.

Sources and research references

Kioxia, imec and Europractice are cited as research references. Their citation does not imply collaboration, affiliation or funding for the project.

Scientific and technological dialogue

Contact MIAR for enquiries about the research, documentation and potential collaborations.

Contact MIAR