Search | arXiv e-print repository

Hypformer: Exploring Efficient Hyperbolic Transformer Fully in Hyperbolic Space

Authors: Menglin Yang, Harshit Verma, Delvin Ce Zhang, Jiahong Liu, Irwin King, Rex Ying

Abstract: Hyperbolic geometry have shown significant potential in modeling complex structured data, particularly those with underlying tree-like and hierarchical structures. Despite the impressive performance of various hyperbolic neural networks across numerous domains, research on adapting the Transformer to hyperbolic space remains limited. Previous attempts have mainly focused on modifying self-attentio… ▽ More Hyperbolic geometry have shown significant potential in modeling complex structured data, particularly those with underlying tree-like and hierarchical structures. Despite the impressive performance of various hyperbolic neural networks across numerous domains, research on adapting the Transformer to hyperbolic space remains limited. Previous attempts have mainly focused on modifying self-attention modules in the Transformer. However, these efforts have fallen short of developing a complete hyperbolic Transformer. This stems primarily from: (i) the absence of well-defined modules in hyperbolic space, including linear transformation layers, LayerNorm layers, activation functions, dropout operations, etc. (ii) the quadratic time complexity of the existing hyperbolic self-attention module w.r.t the number of input tokens, which hinders its scalability. To address these challenges, we propose, Hypformer, a novel hyperbolic Transformer based on the Lorentz model of hyperbolic geometry. In Hypformer, we introduce two foundational blocks that define the essential modules of the Transformer in hyperbolic space. Furthermore, we develop a linear self-attention mechanism in hyperbolic space, enabling hyperbolic Transformer to process billion-scale graph data and long-sequence inputs for the first time. Our experimental results confirm the effectiveness and efficiency of Hypformer across various datasets, demonstrating its potential as an effective and scalable solution for large-scale data representation and large models. △ Less

Submitted 1 July, 2024; originally announced July 2024.

Comments: KDD 2024

arXiv:2402.11352 [pdf, other]

Unified Capacity Results for Free-Space Optical Communication Systems Over Gamma-Gamma Atmospheric Turbulence Channels

Authors: Himani Verma, Kamal Singh

Abstract: Transmit power control, as in the mobile wireless channels, can enable a robust and spectrally efficient communication through atmospheric turbulence in terrestrial free-space optical (FSO) channels. With optical bandwidths in excess of several GHz and eye safety regulations limiting the transmit optical power, the per hertz signal-to-noise ratio (SNR) in terrestrial FSO systems can possibly becom… ▽ More Transmit power control, as in the mobile wireless channels, can enable a robust and spectrally efficient communication through atmospheric turbulence in terrestrial free-space optical (FSO) channels. With optical bandwidths in excess of several GHz and eye safety regulations limiting the transmit optical power, the per hertz signal-to-noise ratio (SNR) in terrestrial FSO systems can possibly become limited, especially true for future high-bandwidth and long-haul applications. Hence, power control becomes significant in terrestrial FSO systems. However, a comprehensive study of dynamic power adaptation in the existing FSO systems is lacking in the literature. In this paper, we investigate FSO communication systems capable of beam power control with heterodyne detection and direct detection based receivers operating under shot noise-limited conditions. Under these considerations, we derive unified exact and asymptotic capacity formulas for the Gamma-Gamma turbulence channels with and without pointing errors; these novel closed-form capacity expressions provide new insights into the impact of varying turbulence conditions and pointing errors. Further, the numerical results highlight the intricate relations of atmospheric turbulence and pointing error parameters in typical terrestrial FSO channel setting. A concrete assessment of the impact of the key channel parameters on the capacity performances of the aforementioned FSO systems is performed revealing several novel and interesting insights. △ Less

Submitted 26 July, 2024; v1 submitted 17 February, 2024; originally announced February 2024.

arXiv:2312.11996 [pdf, other]

Toward Responsible AI Use: Considerations for Sustainability Impact Assessment

Authors: Eva Thelisson, Grzegorz Mika, Quentin Schneiter, Kirtan Padh, Himanshu Verma

Abstract: As AI/ML models, including Large Language Models, continue to scale with massive datasets, so does their consumption of undeniably limited natural resources, and impact on society. In this collaboration between AI, Sustainability, HCI and legal researchers, we aim to enable a transition to sustainable AI development by enabling stakeholders across the AI value chain to assess and quantitfy the env… ▽ More As AI/ML models, including Large Language Models, continue to scale with massive datasets, so does their consumption of undeniably limited natural resources, and impact on society. In this collaboration between AI, Sustainability, HCI and legal researchers, we aim to enable a transition to sustainable AI development by enabling stakeholders across the AI value chain to assess and quantitfy the environmental and societal impact of AI. We present the ESG Digital and Green Index (DGI), which offers a dashboard for assessing a company's performance in achieving sustainability targets. This includes monitoring the efficiency and sustainable use of limited natural resources related to AI technologies (water, electricity, etc). It also addresses the societal and governance challenges related to AI. The DGI creates incentives for companies to align their pathway with the Sustainable Development Goals (SDGs). The value, challenges and limitations of our methodology and findings are discussed in the paper. △ Less

Submitted 19 December, 2023; originally announced December 2023.

arXiv:2310.11182 [pdf, ps, other]

On the Effectiveness of Creating Conversational Agent Personalities Through Prompting

Authors: Heng Gu, Chadha Degachi, Uğur Genç, Senthil Chandrasegaran, Himanshu Verma

Abstract: In this work, we report on the effectiveness of our efforts to tailor the personality and conversational style of a conversational agent based on GPT-3.5 and GPT-4 through prompts. We use three personality dimensions with two levels each to create eight conversational agents archetypes. Ten conversations were collected per chatbot, of ten exchanges each, generating 1600 exchanges across GPT-3.5 an… ▽ More In this work, we report on the effectiveness of our efforts to tailor the personality and conversational style of a conversational agent based on GPT-3.5 and GPT-4 through prompts. We use three personality dimensions with two levels each to create eight conversational agents archetypes. Ten conversations were collected per chatbot, of ten exchanges each, generating 1600 exchanges across GPT-3.5 and GPT-4. Using Linguistic Inquiry and Word Count (LIWC) analysis, we compared the eight agents on language elements including clout, authenticity, and emotion. Four language cues were significantly distinguishing in GPT-3.5, while twelve were distinguishing in GPT-4. With thirteen out of a total nineteen cues in LIWC appearing as significantly distinguishing, our results suggest possible novel prompting approaches may be needed to better suit the creation and evaluation of persistent conversational agent personalities or language styles. △ Less

Submitted 17 October, 2023; originally announced October 2023.

Comments: 6 pages, 1 table, PGAI CIKM 2023

arXiv:2309.10383 [pdf, other]

EdgeP4: A P4-Programmable Edge Intelligent Ethernet Switch for Tactile Cyber-Physical Systems

Authors: Nithish Krishnabharathi Gnani, Joydeep Pal, Deepak Choudhary, Himanshu Verma, Soumya Kanta Rana, Kaushal Mhapsekar, T. V. Prabhakar, Chandramani Singh

Abstract: Tactile Internet based operations, e.g., telesurgery, rely on end-to-end closed loop control for accuracy and corrections. The feedback and control are subject to network latency and loss. We design two edge intelligence algorithms hosted at P4 programmable end switches. These algorithms locally compute and command corrective signals, thereby dispense the feedback signals from traversing the netwo… ▽ More Tactile Internet based operations, e.g., telesurgery, rely on end-to-end closed loop control for accuracy and corrections. The feedback and control are subject to network latency and loss. We design two edge intelligence algorithms hosted at P4 programmable end switches. These algorithms locally compute and command corrective signals, thereby dispense the feedback signals from traversing the network to the other ends and save on control loop latency and network load. We implement these algorithms entirely on data plane on Netronome Agilio SmartNICs using P4. Our first algorithm, $\textit{pose correction}$, is placed at the edge switch connected to an industrial robot gripping a tool. The round trip between transmitting force sensor array readings to the edge switch and receiving correct tip coordinates at the robot is shown to be less than $100~μs$. The second algorithm, $\textit{tremor suppression}$, is placed at the edge switch connected to the human operator. It suppresses physiological tremors of amplitudes smaller than $100~μm$ which not only improves the application's performance but also reduces the network load up to $99.9\%$. Our solution allows edge intelligence modules to seamlessly switch between the algorithms based on the tasks being executed at the end hosts. △ Less

Submitted 19 September, 2023; originally announced September 2023.

arXiv:2305.19120 [pdf, other]

Comparing and combining some popular NER approaches on Biomedical tasks

Authors: Harsh Verma, Sabine Bergler, Narjesossadat Tahaei

Abstract: We compare three simple and popular approaches for NER: 1) SEQ (sequence-labeling with a linear token classifier) 2) SeqCRF (sequence-labeling with Conditional Random Fields), and 3) SpanPred (span-prediction with boundary token embeddings). We compare the approaches on 4 biomedical NER tasks: GENIA, NCBI-Disease, LivingNER (Spanish), and SocialDisNER (Spanish). The SpanPred model demonstrates sta… ▽ More We compare three simple and popular approaches for NER: 1) SEQ (sequence-labeling with a linear token classifier) 2) SeqCRF (sequence-labeling with Conditional Random Fields), and 3) SpanPred (span-prediction with boundary token embeddings). We compare the approaches on 4 biomedical NER tasks: GENIA, NCBI-Disease, LivingNER (Spanish), and SocialDisNER (Spanish). The SpanPred model demonstrates state-of-the-art performance on LivingNER and SocialDisNER, improving F1 by 1.3 and 0.6 F1 respectively. The SeqCRF model also demonstrates state-of-the-art performance on LivingNER and SocialDisNER, improving F1 by 0.2 F1 and 0.7 respectively. The SEQ model is competitive with the state-of-the-art on the LivingNER dataset. We explore some simple ways of combining the three approaches. We find that majority voting consistently gives high precision and high F1 across all 4 datasets. Lastly, we implement a system that learns to combine the predictions of SEQ and SpanPred, generating systems that consistently give high recall and high F1 across all 4 datasets. On the GENIA dataset, we find that our learned combiner system significantly boosts F1(+1.2) and recall(+2.1) over the systems being combined. We release all the well-documented code necessary to reproduce all systems at https://github.com/flyingmothman/bionlp. △ Less

Submitted 30 May, 2023; originally announced May 2023.

Comments: Accepted to the ACL BioNLP Workshop, 2023

arXiv:2305.03845 [pdf, other]

CLaC at SemEval-2023 Task 2: Comparing Span-Prediction and Sequence-Labeling approaches for NER

Authors: Harsh Verma, Sabine Bergler

Abstract: This paper summarizes the CLaC submission for the MultiCoNER 2 task which concerns the recognition of complex, fine-grained named entities. We compare two popular approaches for NER, namely Sequence Labeling and Span Prediction. We find that our best Span Prediction system performs slightly better than our best Sequence Labeling system on test data. Moreover, we find that using the larger version… ▽ More This paper summarizes the CLaC submission for the MultiCoNER 2 task which concerns the recognition of complex, fine-grained named entities. We compare two popular approaches for NER, namely Sequence Labeling and Span Prediction. We find that our best Span Prediction system performs slightly better than our best Sequence Labeling system on test data. Moreover, we find that using the larger version of XLM RoBERTa significantly improves performance. Post-competition experiments show that Span Prediction and Sequence Labeling approaches improve when they use special input tokens (<s> and </s>) of XLM-RoBERTa. The code for training all models, preprocessing, and post-processing is available at https://github.com/harshshredding/semeval2023-multiconer-paper. △ Less

Submitted 5 May, 2023; originally announced May 2023.

Comments: Accepted at the ACL SemEval-2023 Workshop

arXiv:2209.03528 [pdf, ps, other]

CLaCLab at SocialDisNER: Using Medical Gazetteers for Named-Entity Recognition of Disease Mentions in Spanish Tweets

Authors: Harsh Verma, Parsa Bagherzadeh, Sabine Bergler

Abstract: This paper summarizes the CLaC submission for SMM4H 2022 Task 10 which concerns the recognition of diseases mentioned in Spanish tweets. Before classifying each token, we encode each token with a transformer encoder using features from Multilingual RoBERTa Large, UMLS gazetteer, and DISTEMIST gazetteer, among others. We obtain a strict F1 score of 0.869, with competition mean of 0.675, standard de… ▽ More This paper summarizes the CLaC submission for SMM4H 2022 Task 10 which concerns the recognition of diseases mentioned in Spanish tweets. Before classifying each token, we encode each token with a transformer encoder using features from Multilingual RoBERTa Large, UMLS gazetteer, and DISTEMIST gazetteer, among others. We obtain a strict F1 score of 0.869, with competition mean of 0.675, standard deviation of 0.245, and median of 0.761. △ Less

Submitted 12 September, 2022; v1 submitted 7 September, 2022; originally announced September 2022.

Comments: In Proceedings of the Social Media Mining for Health Applications Workshop at COLING 2022

arXiv:2204.06382 [pdf, ps, other]

Empathy-Centric Design At Scale

Authors: Andrea Mauri, Yen-Chia Hsu, Marco Brambilla, Aisling Ann O'Kane, Ting-Hao 'Kenneth' Huang, Himanshu Verma

Abstract: EmpathiCH aims at bringing together and blend different expertise to develop new research agenda in the context of "Empathy-Centric Design at Scale". The main research question is to investigate how new technologies can contribute to the elicitation of empathy across and within multiple stakeholders at scale; and how empathy can be used to design solutions to societal problems that are not only ef… ▽ More EmpathiCH aims at bringing together and blend different expertise to develop new research agenda in the context of "Empathy-Centric Design at Scale". The main research question is to investigate how new technologies can contribute to the elicitation of empathy across and within multiple stakeholders at scale; and how empathy can be used to design solutions to societal problems that are not only effective but also balanced, inclusive, and aware of their effect on society. Through presentations, participatory sessions, and a living experiment -- where data about the peoples' interactions is collected throughout the event -- we aim to make this workshop the ideal venue to foster collaboration, build networks, and shape the future direction of "Empathy-Centric Design at Scale". △ Less

Submitted 13 April, 2022; originally announced April 2022.

Comments: accepted at Workshops at the 2022 CHI Conference on Human Factors in Computing Systems (CHI 2022)

arXiv:2204.06289 [pdf, other]

COCTEAU: an Empathy-Based Tool for Decision-Making

Authors: Andrea Mauri, Andrea Tocchetti, Lorenzo Corti, Yen-Chia Hsu, Himanshu Verma, Marco Brambilla

Abstract: Traditional approaches to data-informed policymaking are often tailored to specific contexts and lack strong citizen involvement and collaboration, which are required to design sustainable policies. We argue the importance of empathy-based methods in the policymaking domain given the successes in diverse settings, such as healthcare and education. In this paper, we introduce COCTEAU (Co-Creating T… ▽ More Traditional approaches to data-informed policymaking are often tailored to specific contexts and lack strong citizen involvement and collaboration, which are required to design sustainable policies. We argue the importance of empathy-based methods in the policymaking domain given the successes in diverse settings, such as healthcare and education. In this paper, we introduce COCTEAU (Co-Creating The European Union), a novel framework built on the combination of empathy and gamification to create a tool aimed at strengthening interactions between citizens and policy-makers. We describe our design process and our concrete implementation, which has already undergone preliminary assessments with different stakeholders. Moreover, we briefly report pilot results from the assessment. Finally, we describe the structure and goals of our demonstration regarding the newfound formats and organizational aspects of academic conferences. △ Less

Submitted 13 April, 2022; originally announced April 2022.

Comments: accepted at Posters and Demos Track at The Web Conference 2022 (WWW 2022)

arXiv:2110.02007 [pdf, other]

doi 10.1016/j.patter.2022.100449

Empowering Local Communities Using Artificial Intelligence

Authors: Yen-Chia Hsu, Ting-Hao 'Kenneth' Huang, Himanshu Verma, Andrea Mauri, Illah Nourbakhsh, Alessandro Bozzon

Abstract: Artificial Intelligence (AI) is increasingly used to analyze large amounts of data in various practices, such as object recognition. We are specifically interested in using AI-powered systems to engage local communities in developing plans or solutions for pressing societal and environmental concerns. Such local contexts often involve multiple stakeholders with different and even contradictory age… ▽ More Artificial Intelligence (AI) is increasingly used to analyze large amounts of data in various practices, such as object recognition. We are specifically interested in using AI-powered systems to engage local communities in developing plans or solutions for pressing societal and environmental concerns. Such local contexts often involve multiple stakeholders with different and even contradictory agendas, resulting in mismatched expectations of these systems' behaviors and desired outcomes. There is a need to investigate if AI models and pipelines can work as expected in different contexts through co-creation and field deployment. Based on case studies in co-creating AI-powered systems with local people, we explain challenges that require more attention and provide viable paths to bridge AI research with citizen needs. We advocate for developing new collaboration approaches and mindsets that are needed to co-create AI-powered systems in multi-stakeholder contexts to address local concerns. △ Less

Submitted 26 April, 2022; v1 submitted 5 October, 2021; originally announced October 2021.

Comments: This manuscript is peer-reviewed and accepted by the Patterns journal

arXiv:2108.13823 [pdf, other]

Temporal Deep Learning Architecture for Prediction of COVID-19 Cases in India

Authors: Hanuman Verma, Saurav Mandal, Akshansh Gupta

Abstract: To combat the recent coronavirus disease 2019 (COVID-19), academician and clinician are in search of new approaches to predict the COVID-19 outbreak dynamic trends that may slow down or stop the pandemic. Epidemiological models like Susceptible-Infected-Recovered (SIR) and its variants are helpful to understand the dynamics trend of pandemic that may be used in decision making to optimize possible… ▽ More To combat the recent coronavirus disease 2019 (COVID-19), academician and clinician are in search of new approaches to predict the COVID-19 outbreak dynamic trends that may slow down or stop the pandemic. Epidemiological models like Susceptible-Infected-Recovered (SIR) and its variants are helpful to understand the dynamics trend of pandemic that may be used in decision making to optimize possible controls from the infectious disease. But these epidemiological models based on mathematical assumptions may not predict the real pandemic situation. Recently the new machine learning approaches are being used to understand the dynamic trend of COVID-19 spread. In this paper, we designed the recurrent and convolutional neural network models: vanilla LSTM, stacked LSTM, ED-LSTM, Bi-LSTM, CNN, and hybrid CNN+LSTM model to capture the complex trend of COVID-19 outbreak and perform the forecasting of COVID-19 daily confirmed cases of 7, 14, 21 days for India and its four most affected states (Maharashtra, Kerala, Karnataka, and Tamil Nadu). The root mean square error (RMSE) and mean absolute percentage error (MAPE) evaluation metric are computed on the testing data to demonstrate the relative performance of these models. The results show that the stacked LSTM and hybrid CNN+LSTM models perform best relative to other models. △ Less

Submitted 31 August, 2021; originally announced August 2021.

Comments: 13 pages

arXiv:2009.14374 [pdf, other]

Rethinking Evaluation Methodology for Audio-to-Score Alignment

Authors: John Thickstun, Jennifer Brennan, Harsh Verma

Abstract: This paper offers a precise, formal definition of an audio-to-score alignment. While the concept of an alignment is intuitively grasped, this precision affords us new insight into the evaluation of audio-to-score alignment algorithms. Motivated by these insights, we introduce new evaluation metrics for audio-to-score alignment. Using an alignment evaluation dataset derived from pairs of KernScores… ▽ More This paper offers a precise, formal definition of an audio-to-score alignment. While the concept of an alignment is intuitively grasped, this precision affords us new insight into the evaluation of audio-to-score alignment algorithms. Motivated by these insights, we introduce new evaluation metrics for audio-to-score alignment. Using an alignment evaluation dataset derived from pairs of KernScores and MAESTRO performances, we study the behavior of our new metrics and the standard metrics on several classical alignment algorithms. △ Less

Submitted 29 September, 2020; originally announced September 2020.

Comments: 10 pages, 6 figures

arXiv:2008.10450 [pdf]

Analysis of COVID-19 cases in India through Machine Learning: A Study of Intervention

Authors: Hanuman Verma, Akshansh Gupta, Utkarsh Niranjan

Abstract: To combat the coronavirus disease 2019 (COVID-19) pandemic, the world has vaccination, plasma therapy, herd immunity, and epidemiological interventions as few possible options. The COVID-19 vaccine development is underway and it may take a significant amount of time to develop the vaccine and after development, it will take time to vaccinate the entire population, and plasma therapy has some limit… ▽ More To combat the coronavirus disease 2019 (COVID-19) pandemic, the world has vaccination, plasma therapy, herd immunity, and epidemiological interventions as few possible options. The COVID-19 vaccine development is underway and it may take a significant amount of time to develop the vaccine and after development, it will take time to vaccinate the entire population, and plasma therapy has some limitations. Herd immunity can be a plausible option to fight COVID-19 for small countries. But for a country with huge population like India, herd immunity is not a plausible option, because to acquire herd immunity approximately 67% of the population has to be recovered from COVID-19 infection, which will put an extra burden on medical system of the country and will result in a huge loss of human life. Thus epidemiological interventions (complete lockdown, partial lockdown, quarantine, isolation, social distancing, etc.) are some suitable strategies in India to slow down the COVID-19 spread until the vaccine development. In this work, we have suggested the SIR model with intervention, which incorporates the epidemiological interventions in the classical SIR model. To model the effect of the interventions, we have introduced \r{ho} as the intervention parameter. \r{ho} is a cumulative quantity which covers all type of intervention. We have also discussed the supervised machine learning approach to estimate the transmission rate (\b{eta}) for the SIR model with intervention from the prevalence of COVID-19 data in India and some states of India. To validate our model, we present a comparison between the actual and model-predicted number of COVID-19 cases. Using our model, we also present predicted numbers of active and recovered COVID-19 cases till Sept 30, 2020, for entire India and some states of India and also estimate the 95% and 99% confidence interval for the predicted cases. △ Less

Submitted 17 August, 2020; originally announced August 2020.

Comments: 27 Pages, 26 Figures, 3 Tables

arXiv:1911.11737 [pdf, other]

Convolutional Composer Classification

Authors: Harsh Verma, John Thickstun

Abstract: This paper investigates end-to-end learnable models for attributing composers to musical scores. We introduce several pooled, convolutional architectures for this task and draw connections between our approach and classical learning approaches based on global and n-gram features. We evaluate models on a corpus of 2,500 scores from the KernScores collection, authored by a variety of composers spann… ▽ More This paper investigates end-to-end learnable models for attributing composers to musical scores. We introduce several pooled, convolutional architectures for this task and draw connections between our approach and classical learning approaches based on global and n-gram features. We evaluate models on a corpus of 2,500 scores from the KernScores collection, authored by a variety of composers spanning the Renaissance era to the early 20th century. This corpus has substantial overlap with the corpora used in several previous, smaller studies; we compare our results on subsets of the corpus to these previous works. △ Less

Submitted 26 November, 2019; originally announced November 2019.

Comments: 8 pages, published at ISMIR 2019

arXiv:1805.02679 [pdf]

Multichannel Distributed Local Pattern for Content Based Indexing and Retrieval

Authors: Sonakshi Mathur, Mallika Chaudhary, Hemant Verma, Murari Mandal, S. K. Vipparthi, Subrahmanyam Murala

Abstract: A novel color feature descriptor, Multichannel Distributed Local Pattern (MDLP) is proposed in this manuscript. The MDLP combines the salient features of both local binary and local mesh patterns in the neighborhood. The multi-distance information computed by the MDLP aids in robust extraction of the texture arrangement. Further, MDLP features are extracted for each color channel of an image. The… ▽ More A novel color feature descriptor, Multichannel Distributed Local Pattern (MDLP) is proposed in this manuscript. The MDLP combines the salient features of both local binary and local mesh patterns in the neighborhood. The multi-distance information computed by the MDLP aids in robust extraction of the texture arrangement. Further, MDLP features are extracted for each color channel of an image. The retrieval performance of the MDLP is evaluated on the three benchmark datasets for CBIR, namely Corel-5000, Corel-10000 and MIT-Color Vistex respectively. The proposed technique attains substantial improvement as compared to other state-of- the-art feature descriptors in terms of various evaluation parameters such as ARP and ARR on the respective databases. △ Less

Submitted 7 May, 2018; originally announced May 2018.

Comments: Accepted in INDICON-2017

arXiv:1004.3270 [pdf]

Optimized Fuzzy Logic Based Framework for Effort Estimation in Software Development

Authors: Vishal Sharma, Harsh Kumar Verma

Abstract: Software effort estimation at early stages of project development holds great significance for the industry to meet the competitive demands of today's world. Accuracy, reliability and precision in the estimates of effort are quite desirable. The inherent imprecision present in the inputs of the algorithmic models like Constructive Cost Model (COCOMO) yields imprecision in the output, resulting in… ▽ More Software effort estimation at early stages of project development holds great significance for the industry to meet the competitive demands of today's world. Accuracy, reliability and precision in the estimates of effort are quite desirable. The inherent imprecision present in the inputs of the algorithmic models like Constructive Cost Model (COCOMO) yields imprecision in the output, resulting in erroneous effort estimation. Fuzzy logic based cost estimation models are inherently suitable to address the vagueness and imprecision in the inputs, to make reliable and accurate estimates of effort. In this paper, we present an optimized fuzzy logic based framework for software development effort prediction. The said framework tolerates imprecision, incorporates experts knowledge, explains prediction rationale through rules, offers transparency in the prediction system, and could adapt to changing environments with the availability of new data. The traditional cost estimation model COCOMO is extended in the proposed study by incorporating the concept of fuzziness into the measurements of size, mode of development for projects and the cost drivers contributing to the overall development effort. △ Less

Submitted 19 April, 2010; originally announced April 2010.

Comments: International Journal of Computer Science Issues online at http://ijcsi.org/articles/Optimized-Fuzzy-Logic-Based-Framework-for-Effort-Estimation-in-Software-Development.php

Journal ref: IJCSI, Volume 7, Issue 2, March 2010

arXiv:1003.1814 [pdf]

An Analytical Approach to Document Clustering Based on Internal Criterion Function

Authors: Alok Ranjan, Harish Verma, Eatesh Kandpal, Joydip Dhar

Abstract: Fast and high quality document clustering is an important task in organizing information, search engine results obtaining from user query, enhancing web crawling and information retrieval. With the large amount of data available and with a goal of creating good quality clusters, a variety of algorithms have been developed having quality-complexity trade-offs. Among these, some algorithms seek to m… ▽ More Fast and high quality document clustering is an important task in organizing information, search engine results obtaining from user query, enhancing web crawling and information retrieval. With the large amount of data available and with a goal of creating good quality clusters, a variety of algorithms have been developed having quality-complexity trade-offs. Among these, some algorithms seek to minimize the computational complexity using certain criterion functions which are defined for the whole set of clustering solution. In this paper, we are proposing a novel document clustering algorithm based on an internal criterion function. Most commonly used partitioning clustering algorithms (e.g. k-means) have some drawbacks as they suffer from local optimum solutions and creation of empty clusters as a clustering solution. The proposed algorithm usually does not suffer from these problems and converge to a global optimum, its performance enhances with the increase in number of clusters. We have checked our algorithm against three different datasets for four different values of k (required number of clusters). △ Less

Submitted 9 March, 2010; originally announced March 2010.

Comments: Pages IEEE format, International Journal of Computer Science and Information Security, IJCSIS, Vol. 7 No. 2, February 2010, USA. ISSN 1947 5500, http://sites.google.com/site/ijcsis/

arXiv:0909.3554 [pdf]

Robustness of the Digital Image Watermarking Techniques against Brightness and Rotation Attack

Authors: Harsh K Verma, Abhishek Narain Singh, Raman Kumar

Abstract: The recent advent in the field of multimedia proposed a many facilities in transport, transmission and manipulation of data. Along with this advancement of facilities there are larger threats in authentication of data, its licensed use and protection against illegal use of data. A lot of digital image watermarking techniques have been designed and implemented to stop the illegal use of the digit… ▽ More The recent advent in the field of multimedia proposed a many facilities in transport, transmission and manipulation of data. Along with this advancement of facilities there are larger threats in authentication of data, its licensed use and protection against illegal use of data. A lot of digital image watermarking techniques have been designed and implemented to stop the illegal use of the digital multimedia images. This paper compares the robustness of three different watermarking schemes against brightness and rotation attacks. The robustness of the watermarked images has been verified on the parameters of PSNR (Peak Signal to Noise Ratio), RMSE (Root Mean Square Error) and MAE (Mean Absolute Error). △ Less

Submitted 18 September, 2009; originally announced September 2009.

Comments: 5 Pages IEEE format, International Journal of Computer Science and Information Security, IJCSIS 2009, ISSN 1947 5500, Impact factor 0.423

Report number: ISSN 1947 5500

Journal ref: Harsh K Verma, Abhishek Narain Singh, Raman Kumar, International Journal of Computer Science and Information Security, IJCSIS, Vol. 5, No. 1, September 2009, USA

Showing 1–19 of 19 results for author: Verma, H