Spectrally encoded single-pixel machine vision using diffractive networks
License
Copyright © 2021 The Authors, some rights reserved; exclusive licensee American Association for the Advancement of Science. No claim to original U.S. Government Works. Distributed under a Creative Commons Attribution NonCommercial License 4.0 (CC BY-NC).
以降,引用記法を用いる図表および英文は,上記ライセンスに基づいて元論文を引用しています.引用記法を用いる日本語はLLM出力です.
Citation
Jingxi Li et al. ,Spectrally encoded single-pixel machine vision using diffractive networks.Sci. Adv.7,eabd7690(2021).DOI:10.1126/sciadv.abd7690
NotebookLM
この研究論文では、ディープラーニングで訓練された回折ネットワークを利用し、物体の空間情報を光のスペクトルに符号化することで、シングルピクセル分光検出器を用いた新しい機械視覚システムが実証されています。このシステムは、光の回折層を用いて画像分類などの計算タスクを全光学的かつ効率的に実行し、特にテラヘルツ帯域での手書き数字の分類で高い精度を達成しました。さらに、検出されたスペクトル情報から画像を再構成する浅い電子ニューラルネットワークを結合することで、タスク固有の画像非圧縮を実証し、光学系と電子系の連携が全体的な推論精度を大幅に向上させることを示しています。このフレームワークは、低遅延かつ低消費電力で、ピクセル数の少ないセンサーシステムでの機械学習タスクを実現する道を開くものです。
本文
・
・
・
To provide a mitigation strategy for this trade-off, next, we introduced a collaboration framework between the diffractive network and its corresponding reconstruction ANN.
このトレードオフの緩和策を提供するために,次に,回折ネットワークとそれに対応する再構成ANNとの間の協調フレームワークを導入した.
This collaboration is based on the fact that our decoder ANN can faithfully reconstruct the images of the input objects using the spectral encoding present in s, even if the optical classification is incorrect, pointing to a wrong class through max(s).
この協調は,光学的分類が正しくなく,max(s)を通して誤ったクラスを指し示していたとしても,我々のデコーダANNがsに存在するスペクトル符号化を用いて,入力オブジェクトの画像を忠実に再構成できるという事実に基づいている.
We observed that by feeding the decoder ANN’s reconstructed images back to the diffractive network as new inputs, we can have it correct its initial wrong inference (see Fig.4 and fig.S2).
デコーダANNの再構成された画像を新たな入力として回折型ネットワークに戻すことで,回折型ネットワークに最初の誤った推論を修正させることができることが確認された(図4,図S2参照).
Through this collaboration between the diffractive network and its decoder ANN, we improved the overall inference accuracy of a given diffractive network model as summarized in Fig.3C and Table 1.
回折型ネットワークとそのデコーダANNのこのような連携により,図3Cと表1にまとめたように,ある回折型ネットワークモデルの全体的な推論精度が向上した.
For example, for the same, highly efficient diffractive network model that was trained using α = 0.4 and β = 0.2, the blind testing accuracy for handwritten digit classification increased from 84.02 to 91.29% (see Figs.3 C and 5B), demonstrating a substantial improvement through the collaboration between the decoder ANN and the broadband diffractive network.
A close examination of Fig.5 and the provided confusion matrices reveal that the decoder ANN, through its image reconstruction, helped correct 870 misclassifications of the diffractive network, resulting in an overall gain/improvement of 7.27% in the blind inference performance of the optical network.
図5と提供された混同行列を精査すると,デコーダANNは,その画像再構成を通じて,回折ネットワークの870の誤分類を修正するのに役立ち,その結果,光ネットワークのブラインド推論性能が全体として7.27%向上/改善したことが明らかになった.
Similar analyses for the other diffractive network models are also presented in figs. S3 to S5.
他の回折ネットワークモデルについても同様の分析を図S3~S5に示す.
Fig. 4. Illustration of the coupling between the image reconstruction ANN and the diffractive network.
https://scrapbox.io/files/68f89a43fddc0a898f5914b8.png
Four MNIST images of handwritten digits are used here for illustration of the concept.
ここでは,概念の説明のために,手書き数字の4つのMNIST画像を使用する.
Two of the four samples, “0” and “3”, are correctly classified by the diffractive network based on max(s) (top green lines), while the other two, “9” and “5”, are misclassified as “7” and “1”, respectively (top red lines).
4つのサンプルのうち2つ,"0 "と "3 "は回折ネットワークによってmax(s)に基づいて正しく分類され(上の緑線),残りの2つ,"9 "と "5 "はそれぞれ "7 "と "1 "に誤分類される(上の赤線).
Using the same class scores (s) at the output detector of the diffractive network, a shallow decoder ANN digitally reconstructs the images of the input objects.
回折ネットワークの出力検出器で同じクラススコア(s)を使って,浅いデコーダANNが入力オブジェクトの画像をデジタル再構成する.
Next, these images are cycled back to the diffractive optical network as new input images, and the new spectral class scores s′ are inferred accordingly, where all of the four digits are correctly classified through max(s′ ) (bottom green lines).
次に,これらの画像は新しい入力画像として回折光学ネットワークに戻され,それに応じて新しいスペクトル・クラス・スコアs′が推論され,max(s′ )によって4桁すべてが正しく分類される(下の緑線).
Last, these new spectral class scores s′ are used to reconstruct the objects again using the same image reconstruction ANN.
最後に,これらの新しいスペクトルクラススコアs′を用いて,同じ画像再構成ANNを用いて再び物体を再構成する.
The blind testing accuracy of this diffractive network for handwritten digit classification increased from 84.02 to 91.29% using this feedback loop (see Figs.3C and 5B).
手書き数字分類のためのこの回折ネットワークのブラインドテスト精度は,このフィードバックループを使うことで84.02%から91.29%に向上した(図3Cおよび5B参照).
This image reconstruction decoder ANN was trained using the MAE loss and softmax cross-entropy loss (see Eq 2).
この画像再構成デコーダANNはMAE損失とソフトマックスクロスエントロピー損失を用いて学習された(式2参照).