MIAR / AI architectures
Research in progress
Research on AI processors with persistent NAND memory and local computation
The project studies architectures in which AI model weights are stored in non-volatile NAND memory and used by local computation circuits. Its research perspective includes a NAND–CMOS chiplet architecture and integration with 3D NAND memory.
Research objective
Explore a self-contained system capable of running AI models with approximately 7 billion parameters, with a target speed of 20 tokens per second and a target power of 30 watts. The research must establish whether these targets can be reached and under what conditions.
Weights are the numerical values that describe a model’s learned behaviour. A token is a unit of text processed or generated; tokens per second indicate generation speed.
Approach
The studied data path connects persistent weight storage to data transfer and computation. The double buffer provides two temporary memory areas to organise alternating data availability; a controlled MAC unit performs multiplication and accumulation operations.
Data path
- NANDPersistent weights · INT4
- Double bufferData transfer and alternation
- MACMultiply and accumulate · INT8 activations
- ResultsOutput values · INT32
In the demonstrator, computation takes place in the FPGA. The NAND–CMOS chiplet architecture and integration with 3D NAND are research perspectives separate from this result.
Preliminary results
Results obtained in simulation
The demonstrator has been verified in simulation on four 16×256 matrices, using INT4 weights, INT8 activations and INT32 results. These names refer to integer representations with 4, 8 and 32 bits, respectively.
- Data-path tests
- 150
- Successfully completed in simulation.
- Tests with non-uniform matrices
- 16
- Successfully completed in simulation.
- Conditions of the CMOS MAC model
- 27
- Successfully verified in the simulated model.
These checks concern the demonstrator and the CMOS MAC model. They do not yet demonstrate execution of the approximately 7-billion-parameter model or the complete processor’s targets of 20 tokens/s and 30 W.
Research status
Simulation, FPGA logic synthesis and experimental validation are distinct levels of verification. Results at one level do not replace the evidence required at the others.
Checks completed
Simulations
Preliminary results are available for the demonstrator on four matrices and for the CMOS MAC model conditions described on this page.
Findings to be documented
FPGA logic synthesis
Synthesis translates the logic description into an implementation intended for the FPGA. Resources, constraints and timing reports must be documented separately; they are not experimental measurements on the prototype.
Still to be performed
Experimental validation
Building the prototype and taking hardware measurements are the next required checks. The complete processor’s targets have not yet been demonstrated.
Next activities
Moving from simulation to a prototype will allow operation to be checked under actual conditions and measurements to be collected for comparison with the research targets.
Prototype implementation
Build the hardware demonstrator and prepare experimental tests of the NAND, double-buffer, MAC and results data path.
Timing and correctness
Verify transfer and computation timing and measure result correctness on the prototype.
Bandwidth, current and energy
Measure data bandwidth, current draw and energy, keeping experimental results separate from design estimates.
Interface and industrial feasibility
Study the internal NAND–CMOS interface and assess the industrial feasibility of the chiplet architecture and integration with 3D NAND.
Documentation
This section is reserved for research documentation. The technical dossier, diagrams and methodology will be linked here when available for consultation. No downloadable files are currently published.
Technical dossier
Description of the demonstrator, preliminary results, architectural hypotheses and verification limits.
Diagrams
Data path, double-buffer organisation and the relationship between memory and computation units.
Methodology
Matrix configuration, numerical formats, data-path tests and CMOS MAC model conditions.
Sources and research references
Kioxia, imec and Europractice are cited as research references. Their citation does not imply collaboration, affiliation or funding for the project.
Scientific and technological dialogue
Contact MIAR for enquiries about the research, documentation and potential collaborations.