Scheduled Maintenance Notice
Please note that Researcher Profiles will be undergoing scheduled maintenance on Wednesday 7th Oct, from 8:00am to 9:00am. During this time, the Researcher Profiles system will be unavailable. We apologise for any inconvenience and appreciate your understanding.
Select Publications
Journal articles
, 2007, 'Trace-based leakage energy optimisations at link time', Journal of Systems Architecture, 53, pp. 1 - 20, http://dx.doi.org/10.1016/j.sysarc.2006.05.002
, 2006, 'Message from HPSEC workshop co-chairs', Proceedings of the International Conference on Parallel Processing Workshops, pp. xix - xxi, http://dx.doi.org/10.1109/ICPPW.2006.46
, 2006, 'Special Section on Parallel/Distributed Computing and Networking', IEICE Transactions on Information and Systems, E89-D, pp. 387 - 388, http://dx.doi.org/10.1093/ietisy/e89-d.2.387
, 2006, 'A lifetime optimal algorithm for speculative PRE', ACM Transactions on Architecture and Code Optimization, 3, pp. 115 - 155, http://dx.doi.org/10.1145/1138035.1138036
, 2006, 'Partial dead code elimination on predicated code regions', Software: Practice and Experience, 36, pp. 1655 - 1685, http://dx.doi.org/10.1002/spe.739
, 2004, 'Efficient and accurate analytical modeling of whole-program data cache behavior', IEEE Transactions on Computers, 53, pp. 547 - 566_3, http://dx.doi.org/10.1109/tc.2004.1275296
, 2002, 'Eigenvectors-Based Parallelisation of Nested Loops with Affine Dependences', Parallel Algorithms and Applications, pp. 237 - 248
, 2002, 'EIGENVECTORS-BASED PARALLELISATION OF NESTED LOOPS WITH AFFINE DEPENDENCES', Parallel Algorithms and Applications, 17, pp. 227 - 248, http://dx.doi.org/10.1080/01495730108941442
, 2002, 'Space-Time Equations for Non-Unimodular Mappings', International Journal of Computer Mathematics, 79, pp. 555 - 572, http://dx.doi.org/10.1080/00207160210953
, 2002, 'Time-minimal tiling when rise is larger than zero', Parallel Computing, 28, pp. 915 - 939, http://dx.doi.org/10.1016/s0167-8191(02)00098-4
, 2000, 'Generating efficient tiled code for distributed memory machines', Parallel Computing, 26, pp. 1369 - 1410, http://dx.doi.org/10.1016/S0167-8191(00)00040-5
, 1999, 'Partitioning and scheduling loops on NOWs', Computer Communications, 22, pp. 1017 - 1033, http://dx.doi.org/10.1016/s0140-3664(99)00073-0
, 1998, 'Reuse-Driven Tiling for Improving Data Locality', International Journal of Parallel Programming, 26, http://dx.doi.org/10.1023/A:1018734612524
, 1997, 'On tiling as a loop transformation', Parallel Processing Letters, 7, pp. 409 - 424, http://dx.doi.org/10.1142/S0129626497000401
, 1997, 'Communication-Minimal Tiling of Uniform Dependence Loops', Journal of Parallel and Distributed Computin, 42, pp. 42 - 59, http://dx.doi.org/10.1006/jpdc.1997.1310
, 1997, 'On Tiling as a Loop Transformation', Parallel Processing Letters, 07, pp. 409 - 424, http://dx.doi.org/10.1142/S0129626497000401
, 1997, 'Unimodular transformations of non-perfectly nested loops', Parallel Computing, 22, pp. 1621 - 1645, http://dx.doi.org/10.1016/S0167-8191(96)00063-4
, 1996, 'Generalising the unimodular approach to restructure imperfectly nested loops', Parallel Processing Letters, 6, pp. 401 - 414, http://dx.doi.org/10.1142/S0129626496000388
, 1996, 'GENERALISING THE UNIMODULAR APPROACH TO RESTRUCTURE IMPERFECTLY NESTED LOOPS', Parallel Processing Letters, 06, pp. 401 - 414, http://dx.doi.org/10.1142/S0129626496000388
, 1996, 'Transformations of nested loops with non-convex iteration spaces', Parallel Computing, 22, pp. 339 - 368, http://dx.doi.org/10.1016/0167-8191(95)00069-0
, 1995, 'Closed-form mapping conditions for the synthesis of linear processor arrays', Journal of VLSI signal processing systems for signal, image and video technology, 10, pp. 181 - 199, http://dx.doi.org/10.1007/BF02407035
, 1994, 'Automating non-unimodular loop transformations for massive parallelism', Parallel Computing, 20, pp. 711 - 728, http://dx.doi.org/10.1016/0167-8191(94)90002-7
, 1992, 'A systolic array for pyramidal algorithms', Journal of VLSI Signal Processing, 4, pp. 89, http://dx.doi.org/10.1007/BF00930620
, 1992, 'ON THE LOADING, RECOVERY AND ACCESS OF STATIONARY DATA IN SYSTOLIC ARRAYS', LECTURE NOTES IN COMPUTER SCIENCE, 634, pp. 259 - 264, https://www.webofscience.com/api/gateway?GWVersion=2&SrcApp=PARTNER_APP&SrcAuth=LinksAMR&KeyUT=WOS:A1992KQ20400031&DestLinkType=FullRecord&DestApp=ALL_WOS&UsrCustomerID=891bb5ab6ba270e68a29b250adbe88d1
, 1992, 'The synthesis of control signals for one-dimensional systolic arrays', Integration, the VLSI Journal, 14, pp. 1 - 32, http://dx.doi.org/10.1016/0167-9260(92)90008-M
, 1991, 'A systolic array for pyramidal algorithms', Journal of VLSI signal processing systems for signal, image and video technology, 3, pp. 237 - 257, http://dx.doi.org/10.1007/BF00925834
, 1991, 'SPECIFYING CONTROL SIGNALS FOR SYSTOLIC ARRAYS BY UNIFORM RECURRENCE EQUATIONS', Parallel Processing Letters, 01, pp. 83 - 93, http://dx.doi.org/10.1142/S0129626491000033
, 1988, 'A new data structure for representing cell hierarchy in layout design', Computers & Graphics, 12, pp. 341 - 348, http://dx.doi.org/10.1016/0097-8493(88)90055-6
Conference Papers
, 2026, 'Beyond k-Limiting: Pointer-Flow-Guided Context Sensitivity for Scalable and Precise Rust Pointer Analysis', in Leibniz International Proceedings in Informatics Lipics, http://dx.doi.org/10.4230/LIPIcs.ECOOP.2026.1
, 2026, 'Field-Sensitive Over-Tainting Reduction in IFDS Taint Analysis via CFL-Reachability', in Leibniz International Proceedings in Informatics Lipics, http://dx.doi.org/10.4230/LIPIcs.ECOOP.2026.10
, 2026, 'Gopher: Efficient Dynamic Graph Pattern Mining via DAG-Driven Execution', in Eurosys 2026 Proceedings of the 2026 European Conference on Computer Systems, pp. 1722 - 1737, http://dx.doi.org/10.1145/3767295.3769365
, 2026, 'Adaptive Draft Sequence Length: Enhancing Speculative Decoding Throughput on PIM-Enabled Systems', in Proceedings International Symposium on High Performance Computer Architecture, http://dx.doi.org/10.1109/HPCA68181.2026.11408598
, 2026, 'ATLAS: Efficient Dynamic GNN System Through Abstraction-Driven Incremental Execution', in Lecture Notes in Computer Science, pp. 17 - 33, http://dx.doi.org/10.1007/978-981-95-1021-4_2
, 2026, 'BIT: Empowering Binary Analysis through the LLVM Toolchain', in Cgo 2026 Proceedings of the 2026 IEEE ACM International Symposium on Code Generation and Optimization, pp. 390 - 402, http://dx.doi.org/10.1109/CGO68049.2026.11395235
, 2026, 'CeDMA: Enhancing Memory Efficiency of Heterogeneous Accelerator Systems Through Central DMA Controlling', in Lecture Notes in Computer Science, pp. 129 - 144, http://dx.doi.org/10.1007/978-981-95-1021-4_10
, 2026, 'DyPARS: Dynamic-Shape DNN Optimization via Pareto-Aware MCTS for Graph Variants', in Cgo 2026 Proceedings of the 2026 IEEE ACM International Symposium on Code Generation and Optimization, pp. 603 - 616, http://dx.doi.org/10.1109/CGO68049.2026.11395218
, 2026, 'FHEFusion: Enabling Operator Fusion in FHE Compilers for Depth-Efficient DNN Inference', in Cgo 2026 Proceedings of the 2026 IEEE ACM International Symposium on Code Generation and Optimization, pp. 70 - 83, http://dx.doi.org/10.1109/CGO68049.2026.11395213
, 2026, 'Meridian: In-Memory Acceleration for RAG with Document Attention Decomposition', in Proceedings International Symposium on Computer Architecture, pp. 387 - 401, http://dx.doi.org/10.1109/ISCA66397.2026.00041
, 2026, 'PriTran: Privacy-Preserving Inference for Transformer-Based Language Models under Fully Homomorphic Encryption', in Cgo 2026 Proceedings of the 2026 IEEE ACM International Symposium on Code Generation and Optimization, pp. 57 - 69, http://dx.doi.org/10.1109/CGO68049.2026.11395232
, 2026, 'Progressive Low-Precision Approximation of Tensor Operators on GPUs: Enabling Greater Trade-Offs between Performance and Accuracy', in Cgo 2026 Proceedings of the 2026 IEEE ACM International Symposium on Code Generation and Optimization, pp. 670 - 682, http://dx.doi.org/10.1109/CGO68049.2026.11395227
, 2026, 'Pyls: Enabling Python Hardware Synthesis with Dynamic Polymorphism via LCRS Encoding', in Cgo 2026 Proceedings of the 2026 IEEE ACM International Symposium on Code Generation and Optimization, pp. 123 - 135, http://dx.doi.org/10.1109/CGO68049.2026.11394843
, 2026, 'TopServe: Task-Operator Co-scheduling for Efficient Multi-DNN Inference Serving on GPUs', in Lecture Notes in Computer Science, pp. 292 - 305, http://dx.doi.org/10.1007/978-3-031-99857-7_21
, 2025, 'Accelerating Delta Debugging through Probabilistic Monotonicity Assessment', in Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering Ease 2025 Edition Ease 2025, pp. 12 - 22, http://dx.doi.org/10.1145/3756681.3756940
, 2025, 'Diff-MoE: Efficient Batched MoE Inference with Priority-Driven Differential Expert Caching', in Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis Sc 2025, pp. 1951 - 1965, http://dx.doi.org/10.1145/3712285.3759903
, 2025, 'TENSORMD: Accelerating Molecular Dynamics with a High-Performance Machine Learning Interatomic Potential', in Proceedings of the International Conference for High Performance Computing Networking Storage and Analysis Sc 2025, pp. 1631 - 1645, http://dx.doi.org/10.1145/3712285.3759844
, 2025, 'MetaHG: Enhancing HGNN Systems Leveraging Advanced Metapath Graph Abstraction', in Eurosys 2025 Proceedings of the 2025 20th European Conference on Computer Systems, pp. 492 - 506, http://dx.doi.org/10.1145/3689031.3717492
, 2025, 'ReSBM: Region-based Scale and Minimal-Level Bootstrapping Management for FHE via Min-Cut', in International Conference on Architectural Support for Programming Languages and Operating Systems ASPLOS, pp. 924 - 939, http://dx.doi.org/10.1145/3669940.3707276
, 2025, 'ANT-ACE: An FHE Compiler Framework for Automating Neural Network Inference', in Cgo 2025 Proceedings of the 23rd ACM IEEE International Symposium on Code Generation and Optimization, pp. 193 - 208, http://dx.doi.org/10.1145/3696443.3708924
, 2025, 'Qiwu: Exploiting Ciphertext-Level SIMD Parallelism in Homomorphic Encryption Programs', in Cgo 2025 Proceedings of the 23rd ACM IEEE International Symposium on Code Generation and Optimization, pp. 523 - 537, http://dx.doi.org/10.1145/3696443.3708917
, 2025, 'Stack Filtering: Elevating Precision and Efficiency in Rust Pointer Analysis', in Cgo 2025 Proceedings of the 23rd ACM IEEE International Symposium on Code Generation and Optimization, pp. 331 - 346, http://dx.doi.org/10.1145/3696443.3708921