As AI adoption expands, the need for real-time processing, reduced reliance on the cloud, and enhanced data privacy is increasing, leading to a rise in AI inference being performed at the edge. Edge AI processes data locally, eliminating the need to send large amounts of data to the cloud, thus contributing to reduced communication costs. It also contributes to improved privacy and security.
Furthermore, low-latency processing enables even faster real-time decision-making, making it effective in applications such as industrial automation, intelligent image analytics, robotics, and other smart edge devices.
MemryX's Value Proposition
To fully realize the benefits of edge AI, a high-performance, power-efficient AI accelerator is essential.
One promising option that is attracting attention is the MemryX edge AI accelerator, which has the following main features:
1. Achieves both high inference performance and excellent power efficiency.
MemryX offers edge AI accelerators that utilize its proprietary "Near-Memory Dataflow Architecture." Its greatest feature is its ability to combine high inference performance with excellent power efficiency.
Traditional GPU-based systems rely on large-scale data transfers to and from external memory. In contrast, MemryX performs calculations near memory, enabling low-latency and highly efficient AI inference. Especially for edge AI applications, it delivers high inference performance with significantly lower power consumption compared to GPUs. Benchmarks also show performance exceeding comparable products, highlighting its practical processing capabilities that cannot be measured solely by simple TOPS values.
2. Ease of implementing AI models
MemryX does not employ the traditional "model zoo" approach to AI model deployment, which limits users to predefined supported models. Instead, it features the ability to compile and deploy existing models developed using common frameworks such as TensorFlow, PyTorch, and ONNX. Furthermore, it supports highly accurate inference using BF16 activation without requiring manual model retraining, extensive optimization, or quantization.
Furthermore, the open-source SDK allows you to compile and test models with a single click, and it also provides a wealth of Python/C++ APIs and sample applications.
3. Flexible scalability to meet performance requirements
On the hardware side, we offer the MemryX Cascade 100 family, our proprietary edge AI inference accelerator platform. Up to 16 MX3 chips can be configured as a single accelerator. This allows users to scale AI inference performance according to application requirements while maintaining the underlying architecture and software environment. Therefore, the Cascade 100 family is suitable for a wide range of applications, from compact embedded systems and smart devices to multi-camera vision systems and high-performance industrial edge servers.
Furthermore, this platform offers broad system flexibility, including simultaneous execution of multiple AI models, support for ARM, x86, and RISC-V host processors, and compatibility with Linux, Windows, and Android.
Product Lineup and Product Specifications
For edge AI applications, we primarily offer MX3 chips and M.2 modules as AI accelerators.
AI accelerator chip (MX3)
The MX3 is designed for integration into products. AI accelerator-tip
[FC-BGA: 9mm professional 9mm】 is.
The main product specifications are as follows:
|
item |
content |
|
Arithmetic performance |
6 TFLOPS (1GHz) |
|
Accuracy of activation values |
BF16 |
|
Maximum number of supported parameters |
21M (4bit), 10.5M (8bit), 5.25M (16bit) |
|
Supported Platforms (Host, OS) |
x86 / ARM / RISC-V Linux / Windows / Android |
|
interface |
PCIe Gen3, 2-lanes, USB 3.0 |
|
Power consumption (typ) |
0.2 to 3W (depending on workload) |
|
Operating temperature range |
-40°C to 125°C (junction temperature) |
M.2 AI Accelerator Module (Cascade 100M)
The MemryX M.2 module is compact, featuring four MX3 chips in a 22mm x 80mm M.2 2280 M-Key form factor.
It is easy to integrate into systems with M.2 slots, allowing you to easily add high-performance AI inference capabilities to industrial PCs, edge computers, and other embedded systems.
Its compact and low-power design makes it ideal not only for designing new systems but also for deployment on existing edge computing platforms.
The main product specifications are as follows:
|
item |
content |
|
Arithmetic performance |
24 TFLOPS (1GHz) |
|
Accuracy of activation values |
BF16 |
|
Maximum number of supported parameters |
84M (4bit), 42M (8bit), 21M (16bit) |
|
Supported Platforms (Host, OS) |
x86 / ARM / RISC-V Linux / Windows / Android |
|
interface |
PCIe Gen3, 2-lanes |
|
Power consumption (typ) |
1 to 12W (depending on workload) |
|
Operating temperature range |
-40℃ to 85℃ (ambient temperature) |
Simplify development with the MemryX SDK.
The MemryX SDK allows you to efficiently create a software environment for developing and deploying applications on the MemryX AI accelerator. Because you can use existing pre-trained models, compilation and deployment can be done without manual retraining, extensive optimization, or quantization. This simplifies the process from model development to deployment to production.
This SDK includes the MemryX Neural Compiler, runtime software, simulator, and development and performance tools, providing an integrated environment for optimizing the performance and efficiency of AI inference.
Developers also have access to over 500 model examples, application demos, setup guides, and other resources covering a wide range of AI workloads and use cases.
The MemryX Developer Hub brings together these tools, documentation, model resources, and examples in one place to support developers from initial integration to production deployment.
Accelerate AI with MemryX — MemryX Developer Hub
1. Getting Started Guide
This guide provides step-by-step instructions, from preparing the host system and setting up the hardware to installing the MemryX Runtime and various tools, and running the sample application.
2. Model Explorer
The MemryX Model Explorer is a reference resource featuring publicly available AI models that have been tested and benchmarked on the MemryX accelerator. It provides developers with model information, performance results, and compiled DFP files to help accelerate application development and deployment.
Importantly, the Model Explorer does not define or restrict the models that can run on MemryX. Customers can use the MemryX toolchain to compile and deploy their own pre-trained or custom-developed models, including those not listed in the Model Explorer. This allows developers to choose the model best suited to their application without being limited to a predefined model library.
3. Tools
The MemryX SDK includes four key tools to support everything from compiling and evaluating AI models to measuring their performance.
• Neural compiler:
This compiler converts trained AI models into Dataflow Programs (DFPs) that can be executed on MemoryX devices. You can compile the model with a single command.
· simulator:
This simulator allows you to evaluate the behavior and performance of compiled models in advance, even in environments without physical hardware. It streamlines verification in the early stages of development.
· benchmark:
This tool measures inference performance on MemryX devices. It allows you to easily check key performance metrics such as latency and FPS.
・DFP Inspect:
This utility allows you to view model information and various metadata contained in a DFP file.
4. API
We provide APIs for Python and C++, allowing you to easily integrate AI inference capabilities into existing applications. The process from sending input data to retrieving inference results is simple to implement, making it useful for developing a variety of edge AI applications.
5. Demo Project
The MemryX Developer Hub offers over 10 tutorial-style demos that provide step-by-step explanations of the entire process, from model download and compilation to preprocessing, inference, and postprocessing. You can learn while understanding the actual flow of AI application development.
Furthermore, MemryX GitHub offers approximately 50 demo projects, which can be used as reference examples for developing various AI applications.
Application example
MemryX provides a high-performance, low-power AI inference environment for a variety of edge AI applications, including surveillance systems, industrial equipment, and robotics.
Smart Surveillance and Smart Cities
It enables real-time analysis of multiple video streams and AI models for applications such as person and vehicle detection, pedestrian and crowd analysis, traffic monitoring, and other intelligent video analytics.
Its scalable architecture allows it to support a wide range of deployment environments, from single cameras and edge devices to multi-camera surveillance systems.
Example: Parking lot availability management and vehicle detection (parking lot management demo on GitHub)
Industrial Automation
This enables real-time AI inference in industrial applications such as visual inspection, defect and anomaly detection, process monitoring, and quality control.
Its high performance and low power consumption enable direct AI processing at the edge, helping manufacturers improve automation, product quality, and operational efficiency.
Example: Detection of abnormalities in the control board