Site Search

MemryX AI Accelerator (MX3 Chip, M.2 Module)_head

As AI adoption expands, the need for real-time processing, reduced reliance on the cloud, and enhanced data privacy is increasing, leading to a rise in AI inference being performed at the edge. Edge AI processes data locally, eliminating the need to send large amounts of data to the cloud, thus contributing to reduced communication costs. It also contributes to improved privacy and security.

Furthermore, low-latency processing enables even faster real-time decision-making, making it effective in applications such as industrial automation, intelligent image analytics, robotics, and other smart edge devices.

MemryX's Value Proposition

To fully realize the benefits of edge AI, a high-performance, power-efficient AI accelerator is essential.

One promising option that is attracting attention is the MemryX edge AI accelerator, which has the following main features:

1. Achieves both high inference performance and excellent power efficiency.

MemryX offers edge AI​ ​accelerators​ ​that utilize its proprietary "Near-Memory Dataflow Architecture." Its greatest feature is its ability to combine high inference performance with excellent power efficiency.

Traditional GPU-based systems rely on large-scale data transfers to and from external memory. In contrast, MemryX performs calculations near memory, enabling low-latency and highly efficient AI inference. Especially for edge AI applications, it delivers high inference performance with significantly lower power consumption compared to GPUs. Benchmarks also show performance exceeding comparable products, highlighting its practical processing capabilities that cannot be measured solely by simple TOPS values.

2. Ease of implementing AI models

MemryX does not employ the traditional "model zoo" approach to AI model deployment, which limits users to predefined supported models. Instead, it features the ability to compile and deploy existing models developed using common frameworks such as TensorFlow, PyTorch, and ONNX. Furthermore, it supports highly accurate inference using BF16 activation without requiring manual model retraining, extensive optimization, or quantization.

Furthermore, the open-source SDK allows you to compile and test models with a single click, and it also provides a wealth of Python/C++ APIs and sample applications.

3. Flexible scalability to meet performance requirements

On the hardware side, we offer the MemryX Cascade 100 family, our proprietary edge AI inference accelerator platform. Up to 16​ ​MX3 chips can be configured as a single accelerator. This allows users to scale AI inference performance according to application requirements while maintaining the underlying architecture and software environment. Therefore, the Cascade 100 family is suitable for a wide range of applications, from compact embedded systems and smart devices to multi-camera vision systems and high-performance industrial edge servers.

Furthermore, this platform offers broad system flexibility, including simultaneous execution of multiple AI models, support for ARM, x86, and RISC-V host processors, and compatibility with Linux, Windows, and Android.

Product Lineup and Product Specifications

For edge AI applications, we primarily offer MX3 chips and M.2 modules as AI accelerators.

AI accelerator chip (MX3)

MemryX MX3 Chip

The MX3 is designed for integration into products. AI accelerator-tip
[FC-BGA: 9mm professional 9mm】 is.

The main product specifications are as follows:

 item

content

Arithmetic performance

6 TFLOPS (1GHz)

Accuracy of activation values

BF16

Maximum number of supported parameters

21M (4bit), 10.5M (8bit), 5.25M (16bit)

Supported Platforms (Host, OS)

x86 / ARM / RISC-V

Linux / Windows / Android

interface

PCIe Gen3, 2-lanes, USB 3.0

Power consumption (typ)

0.2 to 3W (depending on workload)

Operating temperature range

-40°C to 125°C (junction temperature)

 

M.2 AI Accelerator Module (Cascade 100M)

MemryX M.2 Module

The MemryX M.2 module is compact, featuring four​ ​MX3 chips in a 22mm x 80mm​ ​M.2 2280 M-Key form factor.
It is easy to integrate into systems with M.2 slots, allowing you to easily add high-performance AI inference capabilities to industrial PCs, edge computers, and other embedded systems.
Its compact and low-power design makes it ideal not only for designing new systems but also for deployment on existing edge computing platforms.

The main product specifications are as follows:

 item

content

Arithmetic performance

24 TFLOPS (1GHz)

Accuracy of activation values

BF16

Maximum number of supported parameters

84M (4bit), 42M (8bit), 21M (16bit)

Supported Platforms (Host, OS)

x86 / ARM / RISC-V

Linux / Windows / Android

interface

PCIe Gen3, 2-lanes

Power consumption (typ)

1 to 12W (depending on workload)

Operating temperature range

-40℃ to 85℃ (ambient temperature)

 

Simplify development with the MemryX SDK.

The MemryX SDK allows you to efficiently create a software environment for developing and deploying applications on the MemryX AI accelerator. Because you can use existing pre-trained models, compilation and deployment can be done without manual retraining, extensive optimization, or quantization. This simplifies the process from model development to deployment to production.

This SDK includes the MemryX Neural Compiler, runtime software, simulator, and development and performance tools, providing an integrated environment for optimizing the performance and efficiency of AI inference.

Developers also have access to over 500 model examples, application demos, setup guides, and other resources covering a wide range of AI workloads and use cases.

The MemryX Developer Hub brings together these tools, documentation, model resources, and examples in one place to support developers from initial integration to production deployment.

Accelerate AI with MemryX — MemryX Developer Hub

1. Getting Started Guide

This guide provides step-by-step instructions, from preparing the host system and setting up the hardware to installing the MemryX Runtime and various tools, and running the sample application.

Getting Started — MemryX Developer Hub

2. Model Explorer

The MemryX Model Explorer is a reference resource featuring publicly available AI models that have been tested and benchmarked on the MemryX accelerator. It provides developers with model information, performance results, and compiled DFP files to help accelerate application development and deployment.

Importantly, the Model Explorer does not define or restrict the models that can run on MemryX. Customers can use the MemryX toolchain to compile and deploy their own pre-trained or custom-developed models, including those not listed in the Model Explorer. This allows developers to choose the model best suited to their application without being limited to a predefined model library.

Model eXplorer — MemryX Developer Hub

3. Tools

The MemryX SDK includes four key tools to support everything from compiling and evaluating AI models to measuring their performance.

Tools — MemryX Developer Hub

 

Neural compiler:

This compiler converts trained AI models into Dataflow Programs (DFPs) that can be executed on MemoryX devices. You can compile the model with a single command.

 

· simulator:

This simulator allows you to evaluate the behavior and performance of compiled models in advance, even in environments without physical hardware. It streamlines verification in the early stages of development.

 

· benchmark:

This tool measures inference performance on MemryX devices. It allows you to easily check key performance metrics such as latency and FPS.

 

DFP Inspect

This utility allows you to view model information and various metadata contained in a DFP file.

4. API

We provide APIs for Python and C++, allowing you to easily integrate AI inference capabilities into existing applications. The process from sending input data to retrieving inference results is simple to implement, making it useful for developing a variety of edge AI applications.

5. Demo Project

The MemryX Developer Hub offers over 10 tutorial-style demos that provide step-by-step explanations of the entire process, from model download and compilation to preprocessing, inference, and postprocessing. You can learn while understanding the actual flow of AI application development.

Demo — MemryX Developer Hub

 

Furthermore, MemryX GitHub offers approximately 50 demo projects, which can be used as reference examples for developing various AI applications.

MemryX eXamples ー github

Application example

MemryX provides a high-performance, low-power AI inference environment for a variety of edge AI applications, including surveillance systems, industrial equipment, and robotics.

Smart Surveillance and Smart Cities

It enables real-time analysis of multiple video streams and AI models for applications such as person and vehicle detection, pedestrian and crowd analysis, traffic monitoring, and other intelligent video analytics.
Its scalable architecture allows it to support a wide range of deployment environments, from single cameras and edge devices to multi-camera surveillance systems.
Example: Parking lot availability management and vehicle detection (parking lot management demo on GitHub)

Demonstration_Parking Management Using Yolov8s model

Industrial Automation

This enables real-time AI inference in industrial applications such as visual inspection, defect and anomaly detection, process monitoring, and quality control.
Its high performance and low power consumption enable direct AI processing at the edge, helping manufacturers improve automation, product quality, and operational efficiency.

Example: Detection of abnormalities in the control board

pic4_Real-Time PCB Fault Detection1
If a normal circuit board is detected
pic4_Real-Time PCB Fault Detection2
If a circuit board with foreign objects on it is detected

Contact Us