ByteDance High-Resolution AMT System Achieves SOTA in Piano Note and Pedal Transcription
ByteDance introduces a high-resolution piano transcription system trained by regressing the precise onset and offset times of piano notes and pedals.
AI Technology & Industry Review
ByteDance introduces a high-resolution piano transcription system trained by regressing the precise onset and offset times of piano notes and pedals.
PwC and arXiv jointly announced their partnership yesterday, unveiling a convenient new Code tab on the abstract page of arXiv Machine Learning articles.
ICLR 2021 paper An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale suggests Transformers can outperform top CNNs on CV at scale.
NeurIPS 2020 released its list of accepted papers this week with Google, Stanford, and MIT as the top affiliations.
NVIDIA, Mass General Brigham and 20 global hospitals launch federated learning initiative EXAM to build AI model for COVID-19 patient oxygen need prediction
Google AI researchers developed a sign language detection model for video conferencing applications that can perform real-time identification of a person signing as an active speaker.
LPar, a distributed multi agent platform for large scale industrial deployment of polyglot, diverse and inter-operable agents.
Google AI has announced a new audiovisual speech enhancement feature in YouTube Stories (iOS) that enables creators to make better selfie videos by automatically enhancing their voices and reducing noise.
A team from Google, University of Cambridge, DeepMind, and Alan Turing Institute have proposed a new type of Transformer dubbed Performer, based on a Fast Attention Via positive Orthogonal Random features (FAVOR+) backbone mechanism.
Researchers introduced retrieval-augmented generation - a hybrid, end-to-end differentiable model that combines an information retrieval component with a seq2seq generator.
Google Brain recently introduced a new open-sourced TensorFlow package, TensorFlow Recommenders designed to simplify the process of building, evaluating, and serving sophisticated recommender models.
Facebook AI researchers have open-sourced the new wav2vec 2.0 algorithm for self-supervised language learning.
Top Data Scientists Honored for Advanced Research and Applied Data Science in the Field of Knowledge Discovery in Data and Data Mining
The trimmed-down pQRNN extension to Google AI’s projection attention neural network PRADO compares to BERT on text classification tasks for on-device use.
UIUC, Adobe Research and University of Oregon propose HDMatt, a Deep Learning-based image matting Cross-Patch Context module for high-resolution image inputs.
Augmented Temporal Contrast (ATC), a new unsupervised learning (UL) task for learning visual representations agnostic to rewards and without degrading the control policy.
A group of researchers from Google Research and the University of Oxford have introduced a novel technique that can “retiming” people’s movements in videos.
NumPy is the foundation upon which the scientific Python ecosystem is constructed.
From an augmented view of an image, the researchers trained the online network to predict the target network representation of the same image under a different augmented view.
Monster Mash, a novel AI-powered 3D modelling and animation tool, aims to make these arduous 3D animation processes a whole lot easier.
Facebook AI researchers and engineers just made live video content more accessible by enabling automatic closed captions for Facebook Live and Workplace Live.
Microsoft has released four additional DeepSpeed technologies to enable even faster training times, whether on supercomputers or a single GPU.
This research addresses a well-known phenomenon regarding large batch sizes during training and the generalization gap.
OpenAI researchers introduce GPT-f, an automated prover and proof assistant for the Metamath formalization language.
Novel attention condensers designed to enable the building of low-footprint, highly-efficient deep neural networks for on-device speech recognition on the edge.
Researchers introduce a test covering topics such as elementary mathematics, designed to measure language models’ multitask accuracy.
Researchers introduced a novel flow-based video completion algorithm that compares favourably with the state-of-the-art in the field.
This research proposes an efficient and cost-effective solution for multi-frame video interpolation.
The ReDNA Labs research team has devised a new lightweight approach called IGLOO which allows to deal with sequences up to 25,000 steps long.
Although OpenAI hasn’t yet officially announced the GPT-3 pricing scheme, Branwen’s sneak peek has piqued the interest of the NLP community.
DeepMind unveiled a partnership with Google Maps that has leveraged advanced GNNs to improve ETA accuracy.
“Wav2Lip,” a novel lip-synchronization model that outperforms current approaches by a large margin in both quantitative metrics and human evaluations.
Facebook AI this week released a new high-speed library called Opacus.
AMBERT (A Multigrained BERT) leverages both fine-grained and coarse-grained tokenizations to achieve SOTA performance on English and Chinese language tasks.
The paper outlines a predictive model we’ve developed that has the potential to help significantly reduce wasteful healthcare spending.
Researchers propose reducing the workloads of urban planners by introducing deep learning systems to handle some of their responsibilities.
Intel Labs researchers have proposed a novel method for building a robot called “OpenBot” on just a US$50 budget.
Researchers proposed a new model that is designed to spot deepfakes by looking at subtle visual artifacts.
The 16th European Conference on Computer Vision (ECCV) kicked off on Sunday as a fully online conference. In the Conference Opening Session this morning, the ECCV organizing committee announced the conference’s paper submission stats and Best Paper selections.
Researchers from The Chinese University of Hong Kong, Facebook Reality Labs, and Facebook AI Research have unveiled a state-of-the-art monocular 3D hand motion capture method, FrankMocap, which can estimate both 3D hand and body motions from in-the-wild monocular inputs with faster speed and better accuracy than previous approaches.







































