The focus of our work has been to formulate new theoretical models that address significant hurdles, develop applications around them, and realize selected applications on silicon.
Theory
Learning small models
Model complexity, as measured by the Vapnik-Chervonenkis dimension, is related to model parameters in a subtle way. A single-parameter model may have infinite VC dimension, while a model with many parameters may have small VC dimension.
We derive a smooth and differentiable upper bound on the VC dimension that can be used to learn models with small complexity. This approach, called the Minimal Complexity Machine, has led to variants for classification, regression, ELM-like networks, explicit feature maps, and weight quantization. In deep networks, MCM-based learning can show parameter size reductions of 300X.
Theory
Learning from imbalanced datasets
Skewed datasets pose a core challenge for classification. A crude classifier can appear accurate simply by assigning samples to the majority class.
Twin Support Vector Machines address this by learning two non-parallel hyperplanes, each passing through one class and distant from the other. Twin SVM has been widely adapted and has over 1,900 Google Scholar citations.
Theory
Hashing, outliers, and graphs
Hashing-based methods have been adapted to find outliers, detect local and global anomalies, and understand large graphs. This line of work includes approximate nearest-neighbour search in Hamming space, guided random forests, concept drift detection, and graph coarsening.
Recent graph work includes universal graph coarsening, fairness-aware inference through GNN-to-MLP distillation, scalable single-cell data analysis, trustworthy financial-network modeling, robust and fair graph neural networks, and fair federated graph learning.
Theory and Silicon
Swarm intelligence and distributed routing
EigenAnt provides a theoretical understanding of how distributed systems can find shortest paths. The algorithm shows that shortest-path discovery can be viewed as finding the principal eigenvector in a nonlinear dynamical system.
EigenAnt finds the shortest path regardless of initial conditions and parameter choices, making it robust and simple to implement. A distributed router based on this idea was built on a VLSI chip, one of the largest digital designs taped out from IIT Delhi.
EigenAnt based distributed router, chip and PCB.
Healthcare
Rare cells, cancer mutations, and gene panels
Locality-sensitive hashing was used to build a linear-time algorithm for discovering rare anomalous cells from voluminous single-cell expression data. This remains one of the fast rare-cell detection approaches.
Continuous Representation of Codon Switches is a deep learning method for detecting cancer-related somatic mutations without matched normal samples, with applications in cell-free DNA based assessment of tumor mutation burden. Related work identified an 11 platelet-gene panel for blood-based diagnosis of non-small cell lung cancer.
Applications
Material properties
Predicting material properties is challenging because properties depend on composition, processing methods, environment, and factors that may not be fully discernible. Work in this area includes cement literature extraction, symbolic law discovery, oxide glass properties, Brownian dynamics, materials-science LLM evaluation, and low-complexity prediction of oxide glass properties.
Silicon
AI-driven adaptation in analog circuits
Analog circuits are difficult to design because process variation and non-idealities can change characteristics across schematic, layout, and silicon. A Support Vector Machine based A/D converter treats conversion as a set of classifiers, each producing one bit.
Kernel blocks are implemented with analog transconductance-amplifier-like circuits. Digitally programmable current mirrors allow coefficients to be loaded after fabrication, enabling online calibration of the ADC to compensate for non-idealities and environmental degradation.